Kaggle Community Benchmarking Challenge / Original study

Same Score.
Different Failure.

Fechamento BR tests how language models preserve the meaning of fictional numbers, dates, identifiers and missing values. Every source input, gold answer and original response is inspectable below.

28 authored cases84 original answers12 contrast pairs + 4 controls
What the total score hides: Claude and Gemma both scored 26/28. On the same raw money field 4.700, Claude unnecessarily declared ambiguity despite an explicit locale, while Gemma retained the numeric magnitude but dropped required trailing decimals. These are different errors and imply different engineering interventions.

Raw field
Context
Gold answer
Why this is the expected value (never sent to model)

Unedited model responses

Strict scoring requires the correct status, canonical value and JSON structure. A numerically equivalent amount can still violate the expected string format.