August 6, 2026 · Yunus Emre Vurgun

CSV, JSON, or TOON: Choosing a Format for Reference Data

csv · json · toon · formats · tokens

Reference data ships in many formats, and the choice changes both token cost and parse reliability. Here is the trade-off table we use at YJTOON.

The three formats

  • CSV — smallest, but no types, no nesting, and quoting rules that vary by producer.
  • JSON — universal, typed, nested; the most tokens per byte.
  • TOON — YJTOON's compact format: JSON-compatible nesting with tabular notation for uniform arrays, cutting tokens by about 32% across this site's 185 datasets against the JSON we serve, and by more on table-shaped data.

When CSV wins

Flat, rectangular data headed for spreadsheets or pandas. If every row has the same columns and there is no ambiguity about types, CSV is the honest minimum.

When JSON wins

Nested data, tool-calling interfaces that parse JSON natively, and any consumer with a strict JSON schema. Compatibility is a feature; tokens are the cost.

When TOON wins

Uniform object arrays going into LLM prompts, where tokens are the budget. The grammar is small enough to document in one page, and the savings compound when the same dataset is loaded across many turns.

The rule

Pick the format that matches the consumer's parser, then optimize tokens within that constraint. Serving all three — as YJTOON does via ?format= and static files — removes the choice for consumers entirely.