Duane Forrester Decodes

Duane Forrester Decodes

AI is rewriting how brands get found, cited, and trusted. Duane Forrester breaks down what that means for practitioners and marketing leaders who can't afford to get it wrong. duaneforresterdecodes.substack.com

  1. 6d ago

    Google Went “Not Provided” in 2011 and Blinded Us. ChatGPT Just Shipped Its Version.

    Google Went "Not Provided" in 2011 and Blinded Us. ChatGPT Just Shipped Its Version. The organic attribution you keep asking for structurally isn't coming, and OpenAI's own documentation shows why. Here's what to measure instead. This week I'm digging into the question that keeps surfacing, unprompted, in my reader survey: attribution. The organic, item-level attribution we were trained to expect isn't coming from ChatGPT or the other answer engines, and OpenAI's own documentation shows why. They've built the measurement for advertisers and handed organic operators a robots.txt file. I walk through the three different things people actually mean when they say attribution, the first-party measurement you can run today without anyone's permission, a one-question test for any vendor in the space, and why the deeper problem underneath all of it is a trust and literacy gap that belongs to us, not to a platform. Referenced in this episode: The split I kept pointing at between the visibility switches and the measurement stack lives right in OpenAI's crawler documentation, where you'll find the ads bot sitting alongside the search and training bots. https://developers.openai.com/api/docs/bots And here's the paid half I described, OpenAI's own server-to-server conversions pipeline, the pixel and the events API tied to an ads manager account. That's closed-loop attribution that already exists, just behind an ad account. https://developers.openai.com/ads/conversions-api The optimistic end of that wildly uneven conversion range comes from Microsoft Clarity's study of more than 1,200 publisher sites, where AI-referred visitors converted to sign-ups at roughly eleven times the organic-search rate. https://clarity.microsoft.com/blog/ai-traffic-converts-at-3x-the-rate-of-other-channels-study/ And the sober counterweight, the peer-reviewed study in Marketing Science across 973 sites and twenty billion dollars in revenue, found organic LLM traffic converting below every traditional channel except paid social. Both are true, which is the point. https://pubsonline.informs.org/doi/10.1287/mksc.2025.0489 Finally, if you want the longer argument for why the underlying systems, not the dashboards, are where this literacy has to live, that's the whole spine of my book, The Machine Layer. https://www.amazon.com/Machine-Layer-Visible-Trusted-Search/dp/B0G2WZKM59/ref=sr_1_1 Get full access to Duane Forrester Decodes at duaneforresterdecodes.substack.com/subscribe

  2. Jul 12

    Do the Answer Engines Keep Your Fingerprint, or Do They Start Fresh Every Time?

    This week I'm asking whether the fingerprint your SEO work has been pressing into search for years actually carries into the answer engines, or whether they start fresh every time. We walk through where the record provably persists, which is Google's stack and Bing's, where it enters the pipeline and then goes dark, such as in the Bing-to-ChatGPT path, and the layer nobody outside the lab can prove yet, whether the models keep a native, per-domain record of their own the same way traditional search does. Then a practical sorting of your existing work into what pays forward, what does double duty, and what's genuinely new, plus why recovery isn't one thing but a curve. Links and references When I said Google's AI features are rooted in its core ranking systems, that's straight from Google's own blog post on AI Mode: https://blog.google/products/search/ai-mode-search/ On Google measuring sections of a site independently, here's the site reputation abuse policy clarification I referenced: https://developers.google.com/search/blog/2024/11/site-reputation-abuse When I mentioned IndexNow as the pipe that pushes freshness signals to the index, here's the protocol itself: https://www.indexnow.org/ Web IQ came out at Build in 2026, and here’s the link to their own post explaining what it is and what it does: https://www.microsoft.com/en-us/webiq And if you want the longer argument, my book, The Machine Layer: https://www.amazon.com/Machine-Layer-Visible-Trusted-Search/dp/B0G2WZKM59/ref=sr_1_1 The written version of this episode is available as the full article on this same Substack. Get full access to Duane Forrester Decodes at duaneforresterdecodes.substack.com/subscribe

  3. Jul 5

    The Web Is Eating Itself and Your Metrics Look Fine

    Show Notes This week: the web is filling up with machine-written content just as machine readers are about to become its dominant audience, and the two trends feed each other. Retrieval systems already lean toward AI-written text, which quietly narrows the pool of sources behind AI answers while accuracy still looks fine, a state researchers call deceptively healthy. We walk through the mechanism in plain language, why it can't hold, and four things you can do to be the human, evidence-bearing source that survives the shakeout. Links and references On how much of the web is already machine-written, the Graphite analysis as reported by Axios: https://www.axios.com/2025/10/14/ai-generated-writing-humans On the coming explosion in machine-driven queries, Jordi Ribas of Microsoft, on X: https://x.com/JordiRib1/status/2061866606670581871   On retrieval systems preferring AI-written text, the SIGIR study that named the effect, invisible relevance bias: https://arxiv.org/abs/2311.14084 On what happens as synthetic content accumulates in the pool, the 2026 Web Conference paper on retrieval collapse: https://arxiv.org/abs/2602.16136 On why a system that feeds on its own output degrades over time, the Nature research on model collapse: https://www.nature.com/articles/s41586-024-07566-y On the platforms' stated neutrality about how content is made, Google's own guidance on its AI features: https://developers.google.com/search/docs/fundamentals/ai-optimization-guide And if you want the longer argument, my book, The Machine Layer: https://www.amazon.com/Machine-Layer-Visible-Trusted-Search/dp/B0G2WZKM59/ref=sr_1_1 Get full access to Duane Forrester Decodes at duaneforresterdecodes.substack.com/subscribe

  4. Jun 28

    Microsoft Just Proved a Point About Search Today

    Show Notes Microsoft Just Proved a Point About Search Today Announcing Microsoft Web IQ — Bing Search Blog https://blogs.bing.com/search/June-2026/Announcing-Microsoft-Web-IQ Introducing AI Performance in Bing Webmaster Tools (public preview) — Bing Webmaster Blog https://blogs.bing.com/webmaster/February-2026/Introducing-AI-Performance-in-Bing-Webmaster-Tools-Public-Preview New AI Visibility Insights in Bing Webmaster Tools: Intents, Topics, Citation Share, and Compare — Bing Search Blog https://blogs.bing.com/search/June-2026/New-AI-Visibility-Insights-in-Bing-Webmaster-Tools-Intents-Topics-Citation-Share-Compare Jordi Ribas on Web IQ and how AI agents search — Search Engine Land https://searchengineland.com/microsoft-releases-web-iq-powered-by-bing-but-designed-for-how-ai-agents-search-479194 Microsoft Web IQ gives AI agents Bing grounding APIs — Search Engine Journal https://www.searchenginejournal.com/microsoft-web-iq-gives-ai-agents-bing-grounding-apis/577736/ The evolving role of the index: from ranking pages to supporting answers — Bing Search Blog https://blogs.bing.com/search/May-2026/Evolving-role-of-the-index-From-ranking-pages-to-supporting-answers Rank and AI Citation Aren’t the Same Number — Duane Forrester Decodes https://duaneforresterdecodes.substack.com/p/rank-and-ai-citation-arent-the-same The Machine Layer (book) — Amazon https://www.amazon.com/Machine-Layer-Visible-Trusted-Search/dp/B0G2WZKM59/ref=sr_1_1 Get full access to Duane Forrester Decodes at duaneforresterdecodes.substack.com/subscribe

  5. Jun 21

    81.8% of My “AI Assistant” Traffic Was Fake. The Googlebot Number Was Worse.

    Show Notes: 81.8% of My “AI Assistant” Traffic Was Fake Episode summary Over two weeks, a brand-new website with zero promotion behind it logged thirty-three visits from AI assistants. Only six were real. The rest were lying about who they were, and the Googlebot numbers were worse. In this episode I walk through exactly what I found in my own server logs, how I proved each finding past the point of doubt, and the simple method you can run on your own logs this week to see your real numbers. We cover why a bot’s name is a claim and not an identity, the difference between bots that fetch you for a live answer and bots that crawl you to train tomorrow’s models, the one crawler I had to chase four different ways to nail down, and the one major player you structurally cannot measure at all. What you’ll learn: •     Why the bot names in your analytics are a “claims to be” number, not a real one, and the one check that fixes it. •     The 81.8 percent spoof rate hiding in live AI-assistant traffic, and how the fakes gave themselves away. •     Why Googlebot showed 87 percent impersonation, and why that is an old story, not a new one. •     The difference between retrieval crawlers (today’s visibility) and training crawlers (whether the model knows you tomorrow). •     A repeatable, four-step way to settle any bot you cannot verify on the first pass. •     Why Gemini is the one source you cannot measure by name, and how that rhymes with Google’s old “(not provided)” move. The numbers, at a glance: •     Live AI-assistant fetches: 33 claimed, 6 verified, 27 spoofed. An 81.8 percent spoof rate among the requests that could be checked. •     Googlebot: 799 claimed, 107 verified, 692 spoofed. Roughly 87 percent not Google. •     Most active verified crawlers: Anthropic’s ClaudeBot 166, Googlebot 107, OpenAI’s GPTBot 46, OpenAI’s search crawler 40. •     CCBot (Common Crawl): 20 claimed, 0 verified. Confirmed as impostors across four independent checks. A reminder these are two weeks on one small, new site. The method is the point, not my totals. The published IP-range lists (verify your own logs) These are the first-party files each operator publishes. A request is only legitimate if its source IP falls inside the matching list. Each link goes straight to the source. OpenAI ChatGPT-User (live user fetch): https://openai.com/chatgpt-user.json OAI-SearchBot (search / retrieval): https://openai.com/searchbot.json GPTBot (training): https://openai.com/gptbot.json Anthropic (one file covers all of their bots, including ClaudeBot and Claude-User) https://claude.com/crawling/bots.json Perplexity Perplexity-User (live user fetch): https://www.perplexity.com/perplexity-user.json PerplexityBot (crawler): https://www.perplexity.com/perplexitybot.json Google (note: Google moved these to the /crawling/ipranges/ path in 2026, and the old URLs fail quietly) Common crawlers, including Googlebot: https://developers.google.com/static/crawling/ipranges/common-crawlers.json Special-case crawlers: https://developers.google.com/static/crawling/ipranges/special-crawlers.json User-triggered agents: https://developers.google.com/static/crawling/ipranges/user-triggered-agents.json Common Crawl CCBot: https://index.commoncrawl.org/ccbot.json Verification and proof resources •     Google’s full crawler and fetcher reference, where Google states that Google-Extended has no separate user agent and is a robots.txt control token, not a fetcher, and where Google itself warns the user agent string can be spoofed: https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers •     Google’s guide to verifying that a request really came from Google, including the reverse-DNS method: https://developers.google.com/crawling/docs/crawlers-fetchers/verify-google-requests •     Common Crawl’s public index. Drop in a domain and a recent crawl to check whether your site is actually in the corpus. Use a wildcard, for example yoursite.com/*, so you are not just matching the homepage: https://index.commoncrawl.org/ Run it yourself: the four-step chase When a bot will not verify on the first pass, do not stop at “unknown.” Do this: 1.   Check the published IP list. Is the source address inside the operator’s ranges? 2.   Check reverse DNS. Does the IP resolve back to the operator’s own hostname? 3.   Check the corpus or index where one exists, like Common Crawl’s, to see if you were actually captured. 4.   Run a WHOIS lookup on the raw IP to see who really owns it. Commodity hosting in random countries is your answer. Four angles that agree is proof. One that does not is a thread worth pulling. Try it and tell me. Run this on your own logs and send me two numbers: your demand spoof rate, and your Googlebot one. I suspect the real story is in the spread between them. More on the question of what happens to your content after the fetch: https://www.citationiq.com Follow the show so the next episode finds you. Get full access to Duane Forrester Decodes at duaneforresterdecodes.substack.com/subscribe

    81.8% of My “AI Assistant” Traffic Was Fake. The Googlebot Number Was Worse.
  6. Jun 14

    Rank and AI Citation Aren’t the Same Number

    Resources and references Query length, prompts vs. searches: SimilarWeb data on how much longer AI prompts run than Google queries: https://officechai.com/ai/chatgpt-queries-17x-longer-than-google-searches-6x-longer-than-googles-ai-mode-similarweb-data/ Clickstream analysis on the gap between the typed prompt and the search the model actually fires: https://martech.org/chatgpt-growing-as-a-traffic-referrer-reshaping-search-behavior-report-says/ Study on prompt decomposition, multiple retrieval searches per prompt: https://searchengineland.com/chatgpt-search-prompts-data-463407 Longtail and ranking: Why longtail is about specificity and search volume, not word count: https://www.yotpo.com/blog/long-tail-keywords-guide/ On long, specific phrases being easier to rank for at modest authority: https://www.w3era.com/blog/seo/long-tail-keyword-strategy/ On reading search volume as a starting point, not a verdict: https://www.outrank.so/blog/how-to-find-low-competition-keywords Citation vs. organic overlap: Moz, on most AI Mode citations not appearing in the organic results for the same query: https://thenextweb.com/news/ai-changing-seo-tools ZipTie, on how few cited URLs land in Google's top ten: https://ziptie.dev/blog/how-different-ai-platforms-cite-the-same-source-differently/ Semrush AI Mode study, including heavy Perplexity-Google overlap: https://www.semrush.com/blog/ai-mode-comparison-study/ How input shape moves what gets surfaced: Comparative analysis, AI sourcing shifting with the character of the query: https://arxiv.org/abs/2601.16858 Study on outputs shifting when prompts are rephrased: https://arxiv.org/abs/2509.08919 Book: The Machine Layer: https://www.amazon.com/Machine-Layer-Visible-Trusted-Search/dp/B0G2WZKM59/ref=sr_1_1 Get full access to Duane Forrester Decodes at duaneforresterdecodes.substack.com/subscribe

    Rank and AI Citation Aren’t the Same Number
  7. Jun 7

    AI Search Runs on Two Memory Systems. The Platforms Don’t Use Them the Same Way.

    Referenced in this episode: When the Training Data Cutoff Becomes a Ranking Factor (Duane Forrester Decodes) https://duaneforresterdecodes.substack.com/p/when-the-training-data-cutoff-becomes The companion piece this episode builds on, where I first laid out the parametric-versus-retrieval distinction and what it means for timing. How Perplexity finds and chooses its sources (Search Engine Journal) https://www.searchenginejournal.com/perplexity-ai-interview-explains-how-ai-search-works/565395/ Background on why Perplexity runs a live search on essentially every query rather than answering from memory. Google's AI optimization guidance, and why AI Search is still Search (DemandSphere) https://www.demandsphere.com/blog/google-ai-optimization-guide-ai-search-is-still-search/ Support for the point that AI Overviews and AI Mode are served off the core Search index, not from Gemini's parametric memory. Claude web search tool documentation (Anthropic) https://platform.claude.com/docs/en/agents-and-tools/tool-use/web-search-tool Primary source showing Claude's web search runs as a tool the model invokes only when it decides a question needs it. Manage public web access in Microsoft 365 Copilot (Microsoft Learn) https://learn.microsoft.com/en-us/microsoft-365/copilot/manage-public-web-access The admin control behind the point that, on Copilot, whether retrieval happens at all can be a tenant policy setting. Stop Treating AI Visibility as One Problem (Duane Forrester Decodes) https://duaneforresterdecodes.substack.com/p/stop-treating-ai-visibility-as-one The earlier governed-visibility piece this episode zooms into, treating retrieval as one of three layers to manage. ChatGPT search behavior, clickstream insights (Semrush) https://www.semrush.com/blog/chatgpt-search-insights/ The study behind the stat that ChatGPT's share of search-triggering sessions swung between roughly 15 and 66 percent as models updated. Lost in the Middle: How Language Models Use Long Contexts (arXiv) https://arxiv.org/abs/2307.03172 The foundational research on models using long context unevenly, behind the point that being retrieved isn't the same as being used well. How up to date is ChatGPT, and how knowledge cutoffs work (JustDone) https://justdone.com/blog/ai/how-up-to-date-is-chatgpt Context for the training-cadence point that providers now ship frequent point releases, each carrying its own cutoff. The Machine Layer (Amazon) https://www.amazon.com/Machine-Layer-Visible-Trusted-Search/dp/B0G2WZKM59/ref=sr_1_1 My book, for the longer argument on why visibility, trust, and machine-readability are converging into one problem. Get full access to Duane Forrester Decodes at duaneforresterdecodes.substack.com/subscribe

About

AI is rewriting how brands get found, cited, and trusted. Duane Forrester breaks down what that means for practitioners and marketing leaders who can't afford to get it wrong. duaneforresterdecodes.substack.com