Swapping the Name Did Nothing, But Hedging Moved Every Model Source: https://arxiv.org/abs/2608.13328 Paper was published on August 13, 2026 This episode was AI-generated on August 14, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs. The standard fairness test — swap a man's name for a woman's, see what changes — came back completely empty. But adding a few "maybe"s and a "don't you think?" to the same request got a plainer, more hand-holding draft back from GPT-4, Llama, Mistral, and Gemma alike, and probes locate that decision at layer 5 of 28. If the channel that actually moves the output is the one nobody audits, what exactly are the audits catching? Key Takeaways: - Why the counterfactual name-swap audit — ten most common men's names vs. ten most common women's names, appended as a sign-off — produced no measurable difference on any metric - How the authors kill the obvious 'the model just mirrors your style' explanation: prompts differ by fifteen formality points, but prompt formality explains under four percent of response formality, and longer prompts get shorter answers - Where inside the network the decision happens: register decodes at about ninety-nine percent at layer five of twenty-eight, and patching layers zero through seven produces the biggest output shifts - Why steering the register dial breaks the model — push a little too far and it chants "you, you, you" - The steelman: effect sizes are tiny (about a third of a grade level, word count not significant in eleven of twelve cells), and the hedged stimuli were rated markedly less realistic by the authors' own annotators, 3.35 versus 4.33 - The one dimension the model already refuses to copy — prompts seven to sixty times more polite get responses with statistically identical politeness — and why that makes this a design choice rather than a fact of nature 00:00 - The front desk that ignores your badge: The framing beat: signing a prompt with a gendered name changed nothing, while hedged phrasing changed the draft — and why that matters for the emails, cover letters, and resignation letters people actually run through these tools. 01:12 - The boring explanation that has to die: Finn lays out the null hypothesis as strongly as he can — language models are style-matching next-token predictors, so hedgy prompt in, hedgy prose out — and stakes the episode on whether the paper can break it. 01:52 - Four dials, borrowed from 1973: What 'register' means in sociolinguistics, the four features the paper manipulates — hedges, tag questions, collective reference, expressive adjectives — and why leaning on Robin Lakoff's fifty-year-old typology is both pedigree and a fair place to poke. 03:21 - Scaffolding versus deliverable: How the matched-pair stimuli were built from a bit over four hundred real WildChat workplace requests, and the side-by-side mid-year-review email that shows one condition returning a finished draft and the other returning help getting started. 05:37 - Fifteen points in, four percent out: The two regressions and the mediation check that cap how much mirroring could explain — including the negative length coefficient, where longer prompts get shorter responses, which imitation can't produce. 07:34 - The name swap that moved nothing: The two-by-two design crossing register with a 1990 Census name sign-off, where register effects replicated at full strength and name effects came out indistinguishable from noise on every metric. 09:25 - A live sensor wired to nothing: Probes, activation patching, and steering vectors explained, then the finding: both register and name gender are readable at layer five, but only register is causally wired to the output — and pushing on the steering dial breaks the model's coherence. 12:20 - The abstract outruns its own tables: The steelman critique: effect sizes far smaller than the word 'large' implies, word count not significant in eleven of twelve cells, GPT-4 rewrites the authors' own annotators rated unrealistic, and the fact that the authority metric — the one the harm story needs — didn't move. 14:48 - It can already refuse — for politeness: The finding that survives every objection: prompts seven to sixty times more polite yield responses with statistically identical politeness, which turns the whole thing into a changeable design choice about which parts of your voice get copied — plus the authors' proposed fix and the feedback-loop worry. Recommended Reading: - Dialect prejudice predicts AI decisions about people's character, employability, and criminality: The closest large-scale precedent for this episode's central reframe — that how you phrase something, not who you say you are, is the channel where LLMs quietly sort people. (https://arxiv.org/abs/2403.00742) - Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design: The systematic case that trivial surface changes to a prompt swing model behavior, which is the background fact that makes the paper's 'mirroring can't explain this' regressions worth scrutinizing. (https://arxiv.org/abs/2310.11324) - Locating and Editing Factual Associations in GPT: The paper that popularized the activation-patching method Finn walks through, useful for judging what 'the register call happens by layer five' actually licenses you to claim. (https://arxiv.org/abs/2202.05262)