August 6, 2026 · Yunus Emre Vurgun
CSV, JSON, or TOON: Choosing a Format for Reference Data
Reference data ships in many formats, and the choice changes both token cost and parse reliability. Here is the trade-off table we use at YJTOON.
The three formats
- CSV — smallest, but no types, no nesting, and quoting rules that vary by producer.
- JSON — universal, typed, nested; the most tokens per byte.
- TOON — YJTOON's compact format: JSON-compatible nesting with tabular notation for uniform arrays, cutting tokens by about 32% across this site's 185 datasets against the JSON we serve, and by more on table-shaped data.
When CSV wins
Flat, rectangular data headed for spreadsheets or pandas. If every row has the same columns and there is no ambiguity about types, CSV is the honest minimum.
When JSON wins
Nested data, tool-calling interfaces that parse JSON natively, and any consumer with a strict JSON schema. Compatibility is a feature; tokens are the cost.
When TOON wins
Uniform object arrays going into LLM prompts, where tokens are the budget. The grammar is small enough to document in one page, and the savings compound when the same dataset is loaded across many turns.
The rule
Pick the format that matches the consumer's parser, then optimize tokens within that constraint. Serving all three — as YJTOON does via ?format= and static files — removes the choice for consumers entirely.