Eric Freeburg

E. M. Freeburg

Independent research by E. M. Freeburg. Research, Essays, and thought on questions that recur across culture, character, and public life. Free to read in full.

Episodes

  1. 3d ago

    Glaring Fireball

    A novel enters a modern training corpus through a scan, an OCR pass, and a linearizer. It comes out with headings the novelist never wrote, breaks where the thought did not break, and a chapter that ran forty pages cut into eleven sections. Every word survives. The domain label survives. What does not survive is the arrangement — and the arrangement is what the model trains on. The substitution happens upstream of every control the field exercises over training data. Mixture weights choose how often to sample the converted document; filters choose whether to keep it. Both take the corpus as given, and the corpus was given by an extractor. No reported quantity would change if a different extraction library had been chosen. This proposal sets out three studies to establish whether that matters, ordered so the cheapest can end the inquiry. The first needs no models and no training: run current converters over two document classes and count the structural elements in the output with no counterpart in the source. If converters turn out to be conservative on continuous prose, the concern dissolves. The second reads public checkpoints for a predicted divergence between benchmark performance and long-form likelihood. The third — the only expensive one — is a three-arm matched pre-training run separating faithful conversion from imposed hierarchy, which no existing experiment does. 00:21 — Summary03:40 — The gap05:42 — The hypothesis08:29 — Why the gap has persisted10:40 — What has to be distinguished12:37 — The program17:51 — What would follow, and what would not19:32 — Why now rather than later23:08 — Relation to existing work25:01 — What is already known about the model side26:53 — What this would produce

    Glaring Fireball
  2. Mar 27

    The Last Fingerprint: How Markdown Training Shapes LLM Prose

    The em dash is the most discussed tell of machine-written text, and no mechanistic account of it existed. Separately, everyone had noticed that language models default to markdown. Nobody had connected the two. The argument here is that they are one phenomenon: the em dash is markdown leaking into prose — the smallest surviving unit of a structural orientation acquired from markdown-saturated training data, and amplified afterwards by fine-tuning. The paper lays out that genealogy in five steps, from corpus composition through the dash's dual status as both punctuation and structural joint. Then it tests it. Twelve models from five providers were instructed to suppress markdown. Headers, bullets, and bold vanish; em dashes persist — except in Meta's Llama models, which produce none at all. Rates run from 0.0 per thousand words to 9.1 under active suppression. A three-condition gradient shows that even explicit prohibition of the dash fails to eliminate it in some models, and a base-versus-instruct comparison finds the tendency already present before RLHF. The conclusion reframes em dash frequency as a diagnostic of a particular fine-tuning procedure rather than a stylistic defect. 00:18 — Abstract02:27 — 1. Introduction07:22 — 2. Background and Related Work13:58 — 3. The Em Dash Genealogy24:30 — 4. The Two Discourses27:41 — 5. Empirical Evidence42:04 — 6. Discussion50:58 — 7. Conclusion

    The Last Fingerprint: How Markdown Training Shapes LLM Prose

About

Independent research by E. M. Freeburg. Research, Essays, and thought on questions that recur across culture, character, and public life. Free to read in full.