August 6, 2026 · Yunus Emre Vurgun

Units, Dates, and Encodings: The Quiet Data Pitfalls

units · dates · encodings · data-quality

Most reference-data bugs are not logic bugs. They are unit bugs, date bugs, and encoding bugs — the kind that produce confident, wrong answers.

Units belong in the column name

timeout_ms is unambiguous. timeout is a guess. YJTOON datasets name columns with their units (size[bytes], latency_ms) so a truncated or re-ordered table still carries its meaning.

Timestamps: one format, stated

The catalog uses Unix epoch timestamps in the database (created_at) and ISO-8601 in blog metadata (published_at). Both are stated in the docs. The danger is a service that mixes them; the fix is a convention plus a type or name that says which one is which.

Encodings survive only if declared

Every API response includes charset=utf-8. Every consumer should check it. Data with accented characters — café, Üniversität — parses fine in UTF-8 and turns to garbage under Latin-1.

What consumers should verify

  1. Read the units from the column name, not from context.
  2. Confirm the timestamp format before converting — epoch seconds vs ISO-8601 differ by a factor of a thousand in the wrong direction.
  3. Trust the declared charset, not the byte pattern.

None of these are exotic. They are the three quietest ways a correct dataset becomes a wrong answer.