August 6, 2026 · Yunus Emre Vurgun

Reference Data and Prompt Caching: Structure Matters

prompt-caching · tokens · llm · prompts

Modern LLM providers cache prompt prefixes so repeated tokens cost less than fresh ones. Reference data is the ideal cache resident — if you structure it correctly.

How caching interacts with your prompt

Providers cache the longest stable prefix of a prompt. Everything before the first change is reused. If your reference data sits in that stable prefix, every subsequent request that repeats it pays much less.

The rules for cache-friendly reference data

  1. Stable prefix. Put the system prompt and the reference data first; put the variable question last. Do not shuffle the data between turns.
  2. Stable formatting. Use the same format (JSON, YAML, or TOON) every time. A format change invalidates the prefix even if the data is identical.
  3. Versioned snapshots. Reference data should change by version, not by drift. A cached prefix with stale data is worse than a fresh one — pin the version and refresh deliberately.
  4. Trim to the task. Caching reduces token cost, not context bloat. Load only the datasets the task needs.

Why TOON helps here too

Smaller payloads mean more data fits in the cache window, and fewer total tokens per turn. Token-efficient formats and prompt caching compound: the cached prefix is cheaper to store and cheaper to reuse.