Episode description "Influence the parametric side" sounds like something you assign to someone with a deadline. It isn't. This episode prices what actually sits behind that phrase: years of separate parties describing a company in their own words, most of it produced before language models were a consideration, by functions that report to marketing but never owned the output. Research from ICML, ACL, TACL, and Johns Hopkins on how training corpora become model knowledge, why volume of self-published content doesn't substitute for variety of independent description, and why the same property that makes parametric standing impossible to author is what makes it hard to lose. What gets covered ● Why the cost of parametric standing isn't measured in years, and what it's actually measured in ● The difference between training data and parametric standing, and why conflating them causes most of the confusion in this conversation ● Kandpal and colleagues on document count and accuracy, including why a bigger model won't fix a thin footprint ● Allen-Zhu and Li on why repetition from one source does not produce what variety from many sources produces ● The corpus diffusion problem: the most common domain in C4 accounts for less than five hundredths of one percent of documents ● Why polite crawling and robots.txt mean the busiest sites on the internet are underrepresented in training corpora ● Effective cutoffs versus published cutoffs, and why a model's picture of a company is older than its stated date ● The functions that built parametric standing, from PR and analyst relations to crisis communications and review operations ● Why a company controls the review response but never the review ● Knowledge editing research showing that even researchers with direct parameter access can't cleanly change a single fact ● Why the same property that makes this impossible to author is what makes it durable Research referenced Kandpal and colleagues, on long-tail knowledge and pretraining document counts, ICML: arxiv.org/abs/2211.08411 Mallen and colleagues, on entity popularity and parametric reliability, ACL: arxiv.org/abs/2212.10511 Allen-Zhu and Li, on knowledge extractability and phrasing variation, ICML: arxiv.org/abs/2309.14316 Elazar and colleagues, What's In My Big Data, ICLR: arxiv.org/abs/2310.20707 Dodge and colleagues, documenting the Colossal Clean Crawled Corpus, EMNLP: aclanthology.org/2021.emnlp-main.98 Cheng and colleagues, on effective knowledge cutoffs, Johns Hopkins: arxiv.org/abs/2403.12958 Cohen and colleagues, on ripple effects in knowledge editing, TACL: aclanthology.org/2024.tacl-1.16 Common Crawl published crawl statistics: commoncrawl.github.io/cc-crawl-statistics/plots/domains Related The companion article to this episode, which drew the line between retrieval and parametric memory: Entity Mapping Works on Google. Does Any of It Reach ChatGPT? About Duane Forrester Decodes is published weekly at duaneforresterdecodes.substack.com. The Machine Layer is available on Amazon. Get full access to Duane Forrester Decodes at duaneforresterdecodes.substack.com/subscribe