August 6, 2026 · Yunus Emre Vurgun
Units, Dates, and Encodings: The Quiet Data Pitfalls
Most reference-data bugs are not logic bugs. They are unit bugs, date bugs, and encoding bugs — the kind that produce confident, wrong answers.
Units belong in the column name
timeout_ms is unambiguous. timeout is a guess. YJTOON datasets name columns with their units (size[bytes], latency_ms) so a truncated or re-ordered table still carries its meaning.
Timestamps: one format, stated
The catalog uses Unix epoch timestamps in the database (created_at) and ISO-8601 in blog metadata (published_at). Both are stated in the docs. The danger is a service that mixes them; the fix is a convention plus a type or name that says which one is which.
Encodings survive only if declared
Every API response includes charset=utf-8. Every consumer should check it. Data with accented characters — café, Üniversität — parses fine in UTF-8 and turns to garbage under Latin-1.
What consumers should verify
- Read the units from the column name, not from context.
- Confirm the timestamp format before converting — epoch seconds vs ISO-8601 differ by a factor of a thousand in the wrong direction.
- Trust the declared charset, not the byte pattern.
None of these are exotic. They are the three quietest ways a correct dataset becomes a wrong answer.