Justified Posteriors

Seth Benzell and Andrey Fradkin

Explorations into the economics of AI and innovation. Seth Benzell and Andrey Fradkin discuss academic papers and essays at the intersection of economics and technology. empiricrafting.substack.com

  1. 10. Aug.

    Does GDP Growth Mislead Us About Quality of Life?

    In this episode, we discuss “When GDP Misleads: Inferring Living Standards from the Value of a Statistical Life”, written by Stanford/Anthropic economist Chad Jones and Stanford/Epoch economist Philip Trammell. The word “AI” never appears in it, yet every argument about whether AI will show up in the statistics runs through the question of whether the statistics were ever measuring the right thing. We start with priors. Has real GDP per capita over- or underestimated welfare growth? Andrey’s case for over: China’s real GDP per capita is up over 70 times since the early 1950s, and life is not 70 times better. The paper poses a dilemma, suppose you have two goods (food and string quartets) with different productivity growth. The counterintuitive result in the paper is that inventing the low-productivity good makes people better off while slowing measured GDP growth. The paper then proposes to use the value of a statistical life as a measure of welfare. It is the nominal price for being alive, so its growth rate, deflated by the marginal utility of consumption and pinned down by the intertemporal Euler equation, gives you the growth rate of lifetime utility. Seth likes it, while Andrey explains why dissatisfaction with the Euler equation was in the top three reasons he didn’t become a macroeconomist. The results: welfare growth of 2.3% a year since 1940 — and a decline from 1980 to 2024. We get into how VSL is actually estimated (hedonic wage regressions identified off coal miners and other people with unusual preferences about dying), the 1% discount rate that arrives uncited in a one-sentence paragraph, whether mortality risk is being double-counted, and the robustness table where moving the interest rate by a single percentage point in either direction swings the 1940–2020 welfare gain. Links & References The paper * Charles I. Jones & Philip Trammell, “When GDP Misleads: Inferring Living Standards from the Value of a Statistical Life” — headline results: 2.3%/year welfare growth 1940–2024 (6.9×), negative 1980–2024; robustness range 16.1× to 3× on a ±1pp interest-rate change * Charles I. Jones — Stanford GSB * Philip Trammell — Global Priorities Institute, Oxford * Charles I. Jones, “The AI Dilemma: Growth versus Existential Risk” — the previous Chad Jones paper we covered; bounded utility drives an ever-widening wedge between welfare and GDP The value of a statistical life * US DOT departmental guidance on VSL in economic analysis — the paper cites $13.7 million for 2024 (we said $14.5 million on air; see Corrections) * Dora L. Costa & Matthew E. Kahn, “Changes in the Value of Life, 1940–1980” — Journal of Risk and Uncertainty, 2004; the compensating-wage-differential estimates the paper’s VSL time series rests on. “I don’t think any value of statistical life paper is very good.” Concepts discussed * The intertemporal Euler equation — and Andrey’s objections: it fails for a large enough subset of people, individual Euler equations don’t aggregate into a linear functional form, and it implies Ricardian equivalence, which is “so provably false as to invalidate this entire approach” * New goods and variety growth — the food-and-string-quartets example; why the moment of invention (price falling from infinity to finite) is the hard part of any variety adjustment; and the composite-smartphone problem in quality adjustment * Robert Nozick’s experience machine — the eudaimonia button, and what it would mean for measured welfare * Derek Parfit on personal identity and future Tuesday indifference — is your discount rate even something a social welfare function should respect? * Charles I. Jones & Peter J. Klenow, “Beyond GDP? Welfare across Countries and Time” — AER 2016; * Trammell’s bull case for AI welfare: not more stuff per person, but vastly more beings capable of having utility — a total-utilitarian argument that this paper’s single representative agent can’t represent Previously on Justified Posteriors * How much should we invest in AI safety? — our earlier Chad Jones episode (existential risk vs. growth) Corrections * VSL figure: On the episode we said the US DOT value of a statistical life was $14.5 million for 2024. Jones & Trammell cite $13.7 million for 2024 (DOT guidance; current DOT table also lists $13.7M for 2024 / $14.2M for 2025). The slip doesn’t affect the paper’s growth-rate results. * China GDP multiple: On the episode we said China’s real GDP per capita was up over 50× and as high as 72× since 1952/1962. A cleaner figure: 2025 real GDP per capita was about 82× its 1962 level. Chapters * (00:00) Cold open: would you rather be middle-class today, or the king of China? * (00:27) Intro — the paper, and why an AI podcast is covering a paper that never says “AI” * (02:12) Priors: has GDP per capita been a good proxy for welfare? * (03:47) Everything good is correlated with GDP — until you look closely * (05:29) China, 1952 to today: 72× GDP per capita. Is life 72 times better? * (06:32) Two concerns: diminishing returns, and growth in varieties * (07:08) Over or under? Andrey’s split verdict on China and the US * (07:43) 116% since 1980 — “they already had pinball machines” * (08:51) Seth’s prior: diminishing returns dominate, and why the AI age might flip the sign * (10:09) Haven’t we already had huge variety growth? Podcasts, Prairie Home Companion, and the eudaimonia button * (11:15) Putting numbers on it: 95% and 80% that GDP still overstates * (11:36) How much better is life since 1986? Andrey says 25% * (12:12) The benchmark: life expectancy × log consumption, and 41% since 1980 * (13:25) The paper’s setup: food, string quartets, and a productivity gap * (15:05) Why inventing the new good makes us better off and slows measured growth * (16:37) Quality adjustment, and the composite-smartphone problem * (17:08) The problem that exists even before invention: satiation * (17:41) Varieties vs. abundance: the king of China, at length * (19:01) Trammell’s actual bull case: more beings, more utility * (19:44) The clever idea: the value of a statistical life as a nominal price for being alive * (21:10) $14.5 million — but 14.5 million what? * (21:50) The deflator problem, Weimar Germany, and the marginal utility of consumption * (24:12) “Micro or macro?” — the Euler equation and the conditions it needs * (25:54) Discount rate vs. mortality risk — is something being double-counted? * (27:17) The magic equation, in two equivalent forms * (28:59) “The intertemporal Euler is getting some intertemporal shade” * (29:40) Ricardian equivalence, aggregation, and “I have told Chad this” * (31:04) Seth’s defense: welfare on the left, nominal on the right * (31:49) How VSL is actually measured: Costa & Kahn, and the death risk of coal mining * (33:04) Identification off the highest-risk jobs — and the people such jobs attract * (34:44) “It’s just made up”: the $14.5M highway-safety number * (35:22) The other inputs: a 1% discount rate, uncited, and T-bills plus a convenience yield * (35:53) The results: 2.3% a year since 1940 * (37:19) …and negative since 1980. Life peaked in 1984 * (38:05) 6.9× vs 2.4×: could you convince me life is seven times better than 1940? * (39:27) “For whom?” — the representative agent, the 10th percentile, and changing demographics * (41:19) Should a social welfare function respect your discount rate at all? Parfit and future Tuesday indifference * (42:07) Posteriors * (43:37) The robustness table: one percentage point, 16.1× or 3× * (45:10) Sign-off Get full access to Justified Posteriors at empiricrafting.substack.com/subscribe

  2. 27. Juli

    Is AI Replacing Programmers or Boosting Them?

    Justified Posteriors reads “Writing Code vs. Shipping Code” by Mert Demirer, Leon Musolff, and Liyuan Yang In this week’s episode of Justified Posteriors, we update our beliefs with evidence from an ambitious new paper estimating the impact of AI on software production productivity. Demirer (friend of the show), Musolff, and Yang combine public GitHub records for over 100,000 developers with confidential Microsoft data to trace the effect of distinct generations of AI coding tools — autocomplete, sync agents, and async agents — on the code production hierarchy: lines of code, files, commits, pull requests, projects, and releases. The main empirical finding is attenuation of the effect of AI at each step. Enormous gains of 1000% productivity increases or more at the top of the chain translate into about a 30% increase in shipped releases. The model they use to explain this result is closely connected to Kremer’s O-ring logic, which regular listeners will recognize as from a few episodes back. Seth likes the spirit of the model, but feels it is overcomplicated for this context, for reasons he explains. In addition to discussing the data analysis, Andrey and Seth have a good back-and-forth about what we can conclude from it and extrapolate to the economy more generally. The implication that grabbed Seth’s attention is a sentence in this paper’s abstract. Nested in the summary of the careful empirical exercise is an estimated elasticity of substitution of 0.25 between AI and human effort. Big if true! Seth points out the enormous long-run implications of humans being complements to AI in what seems to be the most AI-friendly of tasks: as AI gets cheaper, the human share of income goes up, wages skyrocket, and ultimately AI boosts jobs instead of taking them. That’s a huge real-world hook. Seth and Andrey discuss whether, and if so how much, we update our beliefs in this direction, with Andrey being careful to point out the difficulties of extrapolating from a partial equilibrium elasticity to long-run macro consequences. This episode is sponsored by Revelio Labs — a great source of labor economics data for academics and firms. Now available on WRDS. Priors → Posteriors Prior 1: Does access to AI coding tools boost lines of code written by more than 100%? * Seth: 95% → 99%. Seth came in confident and left more so. Great to see giant numbers. * Andrey: 80% → 95%. A high yes, hedged because “which developers” and “which tools” do a lot of work in that sentence. Prior 2: Does AI boost economic value by 50% or less of the factor by which it boosts lines of code? * Seth: 95% → 97.5%. I have personally produced a great deal of economically worthless code lately. * Andrey: 85% → 95%. Prior 3: Are AI coding tools a gross complement to human labor? Seth’s answer depends on the level of aggregation:The average normie programmer 20%→20% (unchanged)A human engineering department 33%→40%A software company / open source project 60%→85–90%The economy as a whole 33%→33% (unchanged) Andrey: 75% complement at the sectoral level, and he’d put it as low as the programming department — because right now the code that comes out is not shippable without substantial human input. Posterior: still a complement, mildly supported. He declines, on the record and repeatedly, to extrapolate to the macroeconomy. No fun! References The paper under review * Mert Demirer, Leon Musolff & Liyuan Yang, “Writing Code vs. Shipping Code: Productivity Effects Across Generations of AI Coding Tools,” NBER Working Paper 35275 (May 2026). * The authors’ own summary: “Writing code versus shipping code”, VoxEU, June 2026. Prior work by the same team * Kevin Zheyuan Cui, Mert Demirer, Sonia Jaffe, Leon Musolff, Sida Peng & Tobias Salz, “The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers,” Management Science (2026). 4,867 developers, roughly a 26% increase in completed tasks, larger gains for the less experienced. Related Research and Prior Episodes * METR, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity” — the RCT Andrey refers to, in which 16 experienced developers took 19% longer with AI while believing they were 20% faster. arXiv version. * METR’s own update, “We are Changing our Developer Productivity Experiment Design” (Feb 2026) — developers increasingly refuse to be randomized into working without AI, which biases the estimated speedup downward. * Michael Kremer, “The O-Ring Theory of Economic Development,” Quarterly Journal of Economics 108(3), 1993. Our episode on it: Weak Links, Strong Predictions: Kremer’s O-Ring at 30. * Josh Gans & Avi Goldfarb, “O-Ring Automation,” NBER Working Paper 34639 — and our conversation with Avi: Avi Goldfarb on Prediction Machines, O-Ring Tasks, and How AI is Reshaping Economics. * Aaron Chatterji, Tom Cunningham, David Deming, Zoë Hitzig, Christopher Ong, Carl Shan & Kevin Wadman, “How People Use ChatGPT,” NBER Working Paper 34255 — a reference for Andrey’s point about substitution toward home production. Non-work usage grew to over 70% of messages; computer programming is a small share. * Our episode on the same phenomenon in a different market: The Economics of Book Slop — more books, not obviously more valuable books. More apps, not obviously more downloads. Rhymes. * The universal token budget: Alex Imas — Demand Collapse, Bargaining with Machines, and Behavioral AI Economics and Seb Krier on AGI, the Coasean Singularity, and EDM. Table of Contents * Introduction and Today’s Paper — [00:00] * Priors: [02:58] * The Evidence: Ideal and Possible Experiments— [18:34] * The Evidence: Adoption Events and the Attenuating Waterfall — [24:56] * The Evidence: Model and Simulation Results — [42:56] * Aggregates, App Stores, and Posteriors — [56:00] Transcript Introduction and Today’s Paper [00:00] Seth: The abstract of the paper has the following sentence: blah, blah, blah, based on the results in the model, there’s an estimated elasticity of substitution of 0.25, so high complementarity between AI and human effort, which indicates strong complementarities. Wow. As the people say, big if true. Humans and AI, 0.25 complements. Everybody worried about AI taking all our jobs — wrong. All labor share to 100%. Welcome to the Justified Posteriors podcast, the podcast that updates beliefs about the economics of AI and technology. I’m Seth Benzell, with a 0% productivity impact on my code writing, as measured by my podcast release schedule, coming to you from the Pocono Mountains of eastern Pennsylvania. Andrey: And I’m Andrey Fradkin, coming to you from San Francisco, California. Justified Posteriors is sponsored by the fine folks at Revelio Labs, and please do sign up to our podcast and our Substack whenever you get the chance. Seth: Today we’re talking about a really interesting empirical study investigating the impact of AI tool use — autocomplete, synchronous agents, asynchronous agents — on people’s productivity in writing code. This is that kind of hard empirical data that maybe has the potential to move our beliefs. So I’m cautiously optimistic that I’m going to learn a lot from this one. Andrey: It’s the big question, in many ways. We have these tools. We’re using them. What do we get out of them? Are we really that much more productive? That is the question on everyone’s mind, especially since so many of these tools are very costly. There are people token maxing under the belief that the more tokens that are used, the more valuable the output will be. Seth: People are going broke over the tokens. We discussed a universal token budget when Alex Imas was on — or maybe that was with Seb. What are people getting when they’re actually paying for them? The paper we read to look into this is called “Writing Code vs. Shipping Code: Productivity Effects Across Generations of AI Coding Tools,” from Demirer, Leon Musolff, and Liyuan Yang. So, butchered every name, as is common for us. Andrey: Mert is a co-author of the pod, so very excited to be reading a paper of his. Seth: Despite being friends of the show, no punches pulled. Andrey: We never pull any punches. As listeners may know, now is the time for our priors. So what do we think about this topic before we read the paper? Priors: Three Claims, One Spicy [02:58] Seth: Reading this abstract, it seems to make two pretty narrow claims and then one claim that’s a really big spicy one. The first claim is that AI is really productive for helping you write lines of code. The next, more detailed claim is that the translation function going from lines of code to subsequently more advanced stages of production has an attenuation effect. You write more code, then you get more economic stuff out of that code — but the benefit attenuates. And then finally, using a model and some simulation, the paper goes on to argue that this can tell us something about the degree of complementarity between human coders and AIs. So let’s hit those one by one. First off, let me ask you this prior, Andrey. Do you think access to AI coding tools boosts the lines of code written by programmers by more than 100%? Andrey: Yes. I put my prior at 80%. Seth: 80%? That’s not a “yes, shut up.” Why only 80%? Andrey: That’s a pretty high yes. I’m a good Bayesian. Like any empirical question, there are sub-questions about which developers we’re talking about and which specific agentic tools we’re talking about. I’m sure the answer varies by those. It probably increases lines of code written by non-programmers by a ton. Seth: In percentage terms — from, you know, infinity. Andrey: You start with zero and then you go to something. It’s a pretty big percentage increase. It’s really important what samples are being used. But what’s

  3. 13. Juli

    No AI Jobs Apocalypse (Yet) - and a Debt Problem (Now) | Martha Gimbel (Yale Budget Lab)

    This week we’re joined by Martha Gimbel, executive director and co-founder of The Budget Lab at Yale. Martha has worked just about everywhere economic policy gets made — the Joint Economic Committee, the Obama and Biden Councils of Economic Advisers, and Indeed’s Hiring Lab — and she’s now one of the clearest voices on what the data does (and doesn’t yet) say about AI and the labor market. We start in Washington: what politicians are actually asking about AI, the case for “no regrets” economic policy, and why the unemployment insurance system “is not prepared for someone to sneeze within 50 feet of it” — the tech, the financing, and Mississippi’s $200-a-week maximum benefit. Along the way: why almost no policy “pays for itself” (except funding the IRS), whether one hacker per state plus Claude Code can fix government IT, and the Anthropic finding that agentic coding rewards domain expertise, not coding skill. Then the big empirical question: is AI already taking our jobs? Martha walks through the Budget Lab’s labor-market tracking — no sign of broad macro disruption yet — and why the narrative says otherwise: 1.7 million layoffs in a normal month, CEOs with an incentive to blame AI, and Challenger data attributing seven times more layoffs to AI than to tariffs (”this is implausible”). From there we turn to her Senate testimony and Atlantic essay on the national debt: deficits since 2015 are already costing new mortgage holders about $2,500 a year, and “AI will grow us out of it” is a bet — one the Budget Lab has actually modeled. We close with token taxes, sovereign wealth funds (”the great thing about the government — we can tax it”), what CEA is really like from the inside, and a lightning round that ends in a Red Rising roast. Links & References Martha’s work * Martha Gimbel — The Budget Lab at Yale · budgetlab.yale.edu * Martha’s Atlantic essay on how deficits are raising costs for households. * Testimony before the Senate Finance Subcommittee on Fiscal Responsibility and Economic Growth — “The Fiscal Outlook: 2027–2036” hearing (March 11, 2026): debt held by the public ~99% of GDP in 2025 → 120% by 2036 → 175% by 2056 * Evaluating the Impact of AI on the Labor Market: Current State of Affairs — the Budget Lab’s occupational-mix tracking; no sign of broad AI disruption in the macro data yet * What Might AI Adoption Mean for the Fiscal and Economic Outlook? — the AI-and-the-debt scenarios built on the Karger et al. expert forecasts; the fiscal gains are not a free lunch once you add support for displaced workers * Abhi Gupta, The Impact of Deficits on Costs for Households * The Budget Lab Small Macro Model (BLSMM) — the open, interactive macro model discussed in the R-vs-G section * Long-term Impacts of the One Big Beautiful Bill Act — ~zero growth impact at 10 years, negative at 30 (crowding out) * Coming soon from Martha: a token-tax piece in Tax Notes, and Budget Lab work on AI, capital taxation, and sovereign-wealth-fund economics — stay tuned Concepts, papers & people discussed * Anthropic, “Agentic coding and persistent returns to expertise” — the Claude Code study Martha cites: domain expertise, not coding background, predicts success with AI agents * Challenger, Gray & Christmas layoff announcements — the data attributing ~7× more layoffs to AI than to tariffs in 2025 * Brynjolfsson, Chandar & Chen, “Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of AI” — covered previously on the podcast. * Olivier Blanchard, “Public Debt and Low Interest Rates” — the “if r is low, debt has no fiscal cost.” * Danny Yagan and Neil Mehrotra — the Budget Lab’s designated R-vs-G thinkers * The Windfall Trust — the AI economic scenario-planning (”war gaming”) exercises Andrey asks about * The ROAD to Housing Act — “Congress is doing something obviously good. Fingers crossed.” * Reinhart–Rogoff and the Excel error — Seth’s aside on debt-threshold doom predictions * The Council of Economic Advisers and the Joint Economic Committee — partisan vs. unbiased, and the puppy-distribution test * Trade-offs of the pandemic UI plus-up — the $600 flat add-on existed because state systems literally couldn’t compute 90% wage replacement Sci-fi corner * Becky Chambers, The Long Way to a Small, Angry Planet — Martha’s pick for the future we actually want (and a direct appeal: Becky, the people want to know what’s next) * Adrian Tchaikovsky, Children of Time — “so good”; the one with the spiders * Pierce Brown, Red Rising — “angry Harry Potter”; sorry, Basil Our sponsor * This episode is brought to you by Revelio Labs, providing data products useful for many questions about the economy. Chapters * (00:00) Intro & sponsor * (00:59) What politicians are actually asking about AI — and the case for “no regrets” policy * (02:30) D.C.’s misconceptions: politicians are still learning the tools * (03:22) Scenario planning, war gaming, and automatic stabilizers * (04:43) “The UI system is not prepared for someone to sneeze within 50 feet of it” * (05:54) Three reasons UI is broken: the tech, the financing, and Mississippi’s $200/week max * (08:53) Should we federalize it? The IRS, and why almost nothing “pays for itself” (the 7× rule) * (11:28) Why politicians can’t do the obviously good thing: concentrated pain, diffuse benefits * (12:46) Can one hacker per state + Claude Code fix government IT? DOGE, caves full of paper records, and returns to domain expertise * (16:58) Is AI already affecting the labor market? What the macro data shows (and doesn’t) * (19:05) Why the narrative says otherwise: youth outcomes, wrong numbers, and 1.7 million layoffs in a normal month * (22:22) Challenger data: 7× more layoffs attributed to AI than tariffs — “this is implausible” * (23:16) What is behind the entry-level slowdown? The low-hire, low-fire puzzle * (25:16) Will we know it when we see it? Weavers, export controls, and the Napoleonic Wars * (27:54) “I can’t travel to Earth 2”: the counterfactual problem, even at the Budget Lab * (28:43) The fiscal outlook: debt from ~100% to 170%+ of GDP — and why thresholds aren’t the point * (30:44) Fiscal crisis risk, sweet sweet T-bills, and the Blanchard low-rates argument * (32:05) Why didn’t we issue 100-year bonds at 2%? * (32:44) Crowding out: deficits since 2015 ≈ an extra mortgage payment every year * (33:55) Housing affordability, interest rates, and mortgage lock-in * (37:14) What deficit spending crowds out — and the diapers-bill standard for government spending * (40:25) Fiscal gap accounting vs. talking so D.C. actually understands * (44:26) “AI will grow us out of the debt”: R vs. G and the Budget Lab’s AI fiscal scenarios * (47:05) Who pays when the robots work? Payroll taxes, capital taxation, and AI finding every loophole * (48:33) Is capital more or less elastic in the AI age? Token taxes and the IRS * (52:16) A sovereign wealth fund for AI? “The government doesn’t have to own things to get money from things” * (55:41) Fairness and the nerd’s-nerd case for simplifying the tax code * (56:47) Inside the JEC and CEA: partisan vs. unbiased, and the puppy test * (1:00:34) What working at CEA feels like: second best, third best, fourth best * (1:02:06) The CEA junior staff and their “extremely benevolent and well-reasoned rule” * (1:03:34) Budget Lab vs. Penn Wharton vs. CBO: 30-year horizons and what you’re buying with paid family leave * (1:06:38) Private vs. public data: Indeed, benchmarking, and the shutdown’s three contradictory hiring estimates * (1:08:31) What data do we actually want from the AI labs? * (1:09:41) Lightning round: the biggest bottleneck to AI productivity gains (Hollywood vs. healthcare) * (1:10:53) Gun to your head: pre-distribution or redistribute-after? * (1:12:59) Favorite sci-fi: Becky Chambers and the future we want * (1:14:38) Worst sci-fi takes: Red Rising, Children of Time, and a message for Basil * (1:15:35) Sign-off Justified Posteriors is the podcast that updates its beliefs about the economics of AI and technology, hosted by Andrey Fradkin and Seth Benzell. If we changed your priors, subscribe, share it with a friend, and keep your posteriors justified. Transcript What Politicians Are Asking About AI [00:00 – 04:43] [00:00:04] Seth: Welcome to Justified Posteriors, the podcast that updates beliefs about the economics of AI and technology. I’m Seth Benzell, building, with your support, a podcast community which is hopefully more fiscally sustainable than the federal government. Coming to you from the Pocono Mountains of eastern Pennsylvania. Andrey: And I’m Andrey Fradkin, coming to you from San Francisco, California. We are sponsored by the fine folks at Revelio Labs, providing data products useful for many questions about the economy. And we’re very excited to have Martha Gimbel with us today. Martha is the executive director and co-founder of the Yale Budget Lab, and has worked in an enormous variety of impactful roles related to economic policy. Martha, welcome. Martha: Thank you so much for having me. Andrey: To get started: we know that you talk with a lot of politicians and staffers. What questions are they asking you about AI? [00:01:10] Martha: Some of the questions that politicians, policymakers, and staffers are asking are the same ones everyone is asking, right? What is going to happen? How should we think about the economic impacts of this? What is plausible? What seems unlikely? Politicians — they’re just like us. They have the same questions everyone else does. I think the other thing is that people are really trying to figure out how much of the potential future requires something that is different than what we’ve done before. And I don’t just mean “we know that workforce training hasn

  4. 29. Juni

    Litigating the Pope's AI Encyclical with the Lawyers of Scaling Laws Pod

    In this episode of Justified Posteriors, we host Alan Rozenshtein and Kevin Frazier — the law-professor duo behind Lawfare’s Scaling Laws — to take two of the most-discussed AI policy documents of the spring and subject them to an inquisition. Our disputors are probably not what Pope Leo anticipated: two lawyers, two economists, and probably 3/4ths Jewish. Talk about a crossover episode! First up is Pope Leo XIV’s 42,000-word encyclical (that’s Pope-talk for letter) on artificial intelligence. Magnifica Humanitas: On Safeguarding the Human Person in the Time of Artificial Intelligence lays out 5 principles of Catholic social teaching, and then explains how this should shape Catholicism’s approach to AI. We focus on two in particular. The first is subsidiarity, which Seth summarizes as Catholic federalism, the idea that most decisions should be made at as local a level as possible. We discuss both the economic argument for this, but also what the Pope adds to Hayek: Decentralization not for efficiency’s sake, but a kind of ennoblement, the dignity of deciding things locally. The second is the universal destination of goods, which the encyclical extends to “immaterial goods”. This leads to the positive argument of the Pope - that AI should be undertaken as a communal project with decentralized power and discussion, rather than a technocratic “Tower of Babel” that will lead to ruin and division. Much of our disputation focuses on whether these principles actually resolve the important questions. Is the Pope rightfully cautious about an emerging technology, or was this an opportunity to take a stronger stand on what constitutes AI Sin? Interestingly, the Pope’s strongest stand is against transhumanism, which would be a plausible resolution to the dialectic of “Butlerian Jihad” vs. worship of a new machine god. Then we pick up DeepMind’s “Positive Alignment” paper, and the economists get grumpier. Andrey complains that the paper is vacuous, failing to take a stand on actual practical goals or methods. But it sets us up for a good conversation about several issues: Such as liberalism of fear, a type of anti-utopian liberalism; whether “flourishing” is something you can A/B test towards; and where the ‘constitution’ metaphor behind Constitutional AI works vs. breaks down. We also tease a joint project, “SCOTUS Bench,” a new benchmark for evaluating AIs’ ability to predict appeals court outcomes. Watch this space for more on that soon. Related Links * Scaling Laws — Alan and Kevin’s AI, law, and policy podcast at Lawfare * Alan Rozenshtein on X: @ARozenshtein · Kevin Frazier on X: @KevinTFrazier * Magnifica Humanitas — Pope Leo XIV’s first encyclical, “On Safeguarding the Human Person in the Time of Artificial Intelligence,” in full, straight from the Vatican * “Positive Alignment: Artificial Intelligence for Human Flourishing” — the DeepMind-led paper (Laukkonen, Krier, et al.) arguing alignment should optimize toward flourishing, not just away from harm * Claude’s Constitution — Anthropic’s ~20,000-word statement of Claude’s values and character, released under CC0 * “Claude’s Constitution,” with Amanda Askell — the Scaling Laws interview with the document’s primary author (the one we keep saying we’re jealous of) * The Moral Machine — MIT Media Lab’s crowdsourced trolley-problem experiment: millions of judgments on the grandma-versus-criminals ratio * Meta’s Oversight Board — the “Supreme Court of Facebook,” and Kevin’s cautionary tale in institutional design * Andrew B. Hall — Stanford political economist on deliberative democracy, platform governance, and what went wrong with the Oversight Board * The Anthropic Economic Index — the adoption data behind the “whole countries blacked out” point * Judith Shklar, “The Liberalism of Fear” — the cruelty-first, anti-utopian liberalism Alan invokes against thick conceptions of the good Timestamps (00:00) Intro — two papers, four hosts (01:47) Paper 1: Pope Leo XIV’s encyclical, Magnifica Humanitas (04:00) Subsidiarity, or “Catholic federalism” (12:26) Does the Pope take AI seriously enough? Mind-body dualism and the ex cathedra problem (15:34) The coming religious schism over AI personhood — and the Butlerian jihad (18:06) Transhumanism and the dignity of human limits (20:59) When is using AI a sin? Best-man speeches and eulogies (25:05) The universal destination of goods — is AI access already universal? (33:37) Is AI a centralizing technology? Dignity vs. efficiency (36:37) Freedom vs. control, the labor market, and make-work (41:10) Chess, the centaur era, and living after we’re no longer the best (47:34) Sponsor: Revelio Labs (48:49) Paper 2: DeepMind’s “Positive Alignment” (49:17) The liberalism of fear and thick vs. thin notions of the good (53:53) Is positive alignment an empirical question? A/B-testing flourishing (56:29) What would a useful positive-alignment paper actually do? (58:09) Constitutional AI as a site for public participation (1:00:47) The Moral Machine and trolley problems at scale (1:01:08) Does the “constitution” metaphor hold? Virtue ethics and self-binding (1:10:02) Running every Supreme Court case through the models (1:10:53) Lessons from Meta’s Oversight Board (1:15:09) Wrap-up Justified Posteriors is a reader-supported publication. To receive new posts and support our work, consider becoming a free or paid subscriber. You’re also invited to our Discord community at: https://discord.gg/2r3pExumQ Our sponsor This episode is brought to you by Revelio Labs, the leading provider of labor-economics data, available to academics on WRDS. Transcript: Seth (00:00:00): [upbeat music] Welcome to the Justified Posteriors podcast, the podcast that updates beliefs about the economics of AI and technology. I'm Seth Benzell, always positive and always aligned, coming to you from the Pocono Mountains of eastern Pennsylvania. Andrey (00:00:23): And I'm Andrey Fradkin, coming to you from San Francisco, California. We're sponsored by Revelio Labs, fine purveyors of data products. And we're very excited to have Alan Rozenshtein and Kevin Frazier from the "Scaling Laws" podcast on the podcast today. Welcome. Alan (00:00:43): Yeah, thanks for having us. Andrey (00:00:45): Just for our listeners, why don't you tell us a little bit about "Scaling Laws?" Kevin (00:00:51): Sure thing. So our main goal here is to provide robust and timely analysis of all AI policy questions. And that's an expansive ambit, and it's one that keeps us really, really busy because if it's not an executive order, then it's some big new policy idea from one of the labs, or it's some new economic report. But really what we try to do is dive into the weeds of policy and legal issues that are emerging in the AI space, given our backgrounds as law professors. But Alan's the one with the brain, so I'll let him fill in the details on earth. Alan (00:01:30): No, that's a perfect description. Yeah. We just think that there's a lot of really interesting stuff happening at the intersection of AI, law, policy, especially around national security, which is the core focus of the publication that "Scaling Laws" is part of, which is Lawfare, and so we're trying to fight the good fight, and it's never a dull moment. Kevin (00:01:47): That out of the way, I think we can dive into our first paper. Although, I think by any podcast standards, our first paper is lengthy to say the least, dealing with the- Andrey (00:02:00): Mm Kevin (00:02:01): ... pope's encyclical at 42,000 words, or for all those listening, about two and a half hours on my stationary bike. I don't know what that says about my biking skill- Andrey (00:02:13): [laughing] Kevin (00:02:14): ... or my reading ability, but it was a very tiring afternoon. But a very extensive, very important read from Pope Leo. And this has been covered by a lot of folks, but I don't think it's ever been covered by two lawyers and two economists at once. Andrey (00:02:34): [laughing] Kevin (00:02:36): My hunch is that this wasn't what Pope Leo was anticipating when he was sitting and putting a... I like to think of him writing with a quill and- Andrey (00:02:45): [laughing] Kevin (00:02:45): ... on some very old paper. But, I don't think he anticipated this podcast duo diving into his encyclical. Seth (00:02:55): Well, hopefully, our analysis will be less a Tower of Babel of technocratic overreach, and more a blessed city of Jerusalem built together by our common efforts. Kevin (00:03:06): Seth did his reading. Seth dove in. All right. Excellent. Good to hear. Well, I think at this point in time, we're talking in early June. By the time folks are listening to this, unless you've been living under- Seth (00:03:18): There may be a new encyclical. [laughing] Kevin (00:03:21): Somebody- Seth (00:03:22): Pope maybe changed his mind. Kevin (00:03:24): Yeah. Ugh. But there's so much to cover in this encyclical. Obviously, we could start with just the pope's analysis of the role of the Church and of social doctrine, which he gets into in extensive detail, and that covers about 20 to 30 pages. I think for the sake of our podcast, that's probably not our main forte in terms of analyzing the evolution of the Church's social doctrine. But I will let anyone intervene there if they're extremely fired up about that posture. Andrey (00:04:00): [chuckles] Kevin (00:04:00): But I do think that the first area for us to really explore, that both economists and lawyers can appreciate, is this idea of subsidiarity, which is really- Andrey (00:04:12): Mm Kevin (00:04:12): ... the notion that we have various institutions operating at various levels of jurisdiction, and that ultimately we want to devolve regulation or governance of an issue to the smallest capable actor. And that has a lot of resonance and a lot of power in the Church's teach

  5. 15. Juni

    Ioana Marinescu on Insuring Workers for AI, Monopsony, and Philosophy

    This week we’re joined by Ioana Marinescu, labor economist at the University of Pennsylvania’s School of Social Policy & Practice, former Principal Economist at the U.S. Department of Justice Antitrust Division, and a member of Anthropic’s Economic Advisory Board. Ioana is one of the people who put labor-market monopsony on the antitrust map, and she’s now thinking hard about what the social safety net should look like if AI hits the labor market the way the optimists (and the doomers) say it might. We start with her Digitalist Papers essay, which proposes a flexible, two-tier toolkit: AI Adjustment Insurance (extended unemployment benefits + retraining + wage insurance, modeled on Trade Adjustment Assistance) for the churn scenario, and a scalable Digital Dividend — a broad-based cash transfer funded by a small tax on the digital sector — for the world where the jobs don’t come back. Along the way: whether to make policy now or wait, what counts as the “status quo,” moral hazard in mass unemployment, the TAA wage-insurance result that repaid its own subsidy, and Andrey’s “we can’t afford UBI” pushback. Then we get into her new model with Konrad Kording, (Artificial) Intelligence Saturation and the Future of Work”— why splitting the economy into an intelligence sector and a physical sector implies that output and wages saturate even as AI scales to infinity, the robots-vs-LLMs debate, and whether to just relabel “physical” as the non-automatable sector. We close with her DOJ years: defining monopsony, the transmigrante used-car collusion-and-murder case, the Penguin Random House–Simon & Schuster merger (yes, Stephen King testified), antitrust and AI, and a lightning round on ikigai, Camus, and Rawls vs. Mill. Links & References Ioana’s work * marinescu.eu — Ioana’s website · Penn SP2 faculty page * Ioana Marinescu, “Resilient by Design: Dual Safety Nets for Workers in the AI Economy” — The Digitalist Papers, Vol. 2: The Economics of Transformative AI (volume) * Konrad Kording & Ioana Marinescu, “(Artificial) Intelligence Saturation and the Future of Work” — working paper (Brookings write-up & interactive tool). The model finds wage growth can reverse once roughly a third of intelligence tasks are automated. * Ioana Marinescu, comments on Betsey Stevenson’s chapter — NBER, The Economics of Artificial Intelligence: An Agenda (the ikigai discussion) Concepts, papers & people discussed * Trade Adjustment Assistance (TAA) — the template for Ioana’s adjustment insurance; the wage-insurance component that got people back to work faster and was net fiscally positive * Betsey Stevenson, “Artificial Intelligence, Income, Employment, and Meaning” — the post-AGI meaning / ikigai argument Ioana was commenting on * “GPTs are GPTs” — Eloundou, Manning, Mishkin & Rock, GPTs are GPTs: An Early Look at the Labor Market Impact Potential of LLMs — the occupational LLM-exposure measure (”Eloundou et al. / Daniel Rock”) correlated with COVID-era telework * Pascual Restrepo — job-market work on skill mismatch and structural unemployment during automation waves * Daron Acemoglu & Pascual Restrepo, “Robots and Jobs: Evidence from US Labor Markets”. * Albert Camus, The Myth of Sisyphus; ikigai (Japanese: “reason for being”) * Baumol’s cost disease * John Rawls and John Stuart Mill (Utilitarianism) Antitrust & the DOJ * The DOJ Antitrust Division, monopsony in the labor market, and the 2023 Merger Guidelines * Judge blocks the Penguin Random House–Simon & Schuster merger (2022) on a labor theory of harm to authors — Stephen King testified for the government * The transmigrante used-car export case — collusion (and worse) in the US-to-Latin America used-car trade * Anthropic’s Economic Index and Economic Advisory Board * Leopold Aschenbrenner’s Situational Awareness — the “we’ll have to nationalize it” argument referenced on consolidation Previously on Justified Posteriors * Our episode on the Anthropic Economic Index. Our sponsor * This episode is brought to you by Revelio Labs, the leading provider of labor-economics data, available to academics on WRDS. Chapters * (00:00) Intro & sponsor * (00:47) The Digitalist Papers proposal: a flexible safety net for the AI labor shock — and why make policy now * (03:48) Why unemployment insurance isn’t enough, and the Trade Adjustment Assistance template * (05:51) What counts as the “status quo”? Banning AI vs. letting it run * (07:42) How much to insure: moral hazard, mass unemployment, and the three parts of AI Adjustment Insurance * (11:15) Skill mismatch (Restrepo), and how do you certify a layoff was “due to AI”? * (14:45) Did TAA buy social buy-in for free trade? Underfunding — and the wage insurance that repaid its own subsidy * (16:38) “Would Hillary be president?” General-equilibrium pushback and the ski-instructor problem * (19:28) Will the new jobs still be there in two years? The lump-of-labor fallacy * (22:09) Policy B: the Digital Dividend — unconditional, broad-based cash from a small digital-sector tax * (23:52) How to fund it: a sales tax, a sovereign-style fund, and deliberately slowing diffusion a little * (26:00) “We can’t afford UBI”: productivity growth, 0.5% vs. the deficit, and setting money aside ex ante * (30:47) Taxing digital goods: VPNs, evasion, and land-value taxes * (34:23) The motte-and-bailey worry, and the other reasons to like UBI * (36:05) The new model: (Artificial) Intelligence Saturation — intelligence vs. physical sectors, and the telework × AI-exposure correlation * (40:14) Gross complements: why output and wages saturate even with infinite intelligence * (42:23) Won’t enough intelligence just automate the physical world? Robots vs. LLMs * (45:52) “15% by 2030”: humanoid robots, cost, and bespoke vs. general-purpose machines * (47:58) Baumol, the “humanness sector,” and relabeling physical as the non-automatable sector * (48:52) The capital-share / profit-share puzzle: if they’re complements, why has the intelligence share risen? * (50:25) The DOJ years: monopsony, and what the Antitrust Division actually does (mid-roll sponsor at 51:29) * (54:52) “Assassinating rival CEOs”: the transmigrante collusion-and-murder case * (58:12) Favorite cases: Stephen King, the publisher merger, and the chicken-farmer monopsony settlement * (1:01:30) Antitrust and AI: foundation models, consolidation, and the natural-monopoly question * (1:06:05) Slowing AI by allowing market power; Leopold, nationalization, and diminishing returns vs. the singularity * (1:09:27) Substitutability, the AK economy, and short-run vs. long-run wages * (1:10:59) Lightning round: ikigai, Camus, and the myth of Sisyphus * (1:12:44) Can we build market-like mechanisms for ikigai? Loneliness and coordination costs * (1:14:13) The Anthropic Economic Advisory Board and the Economic Index * (1:15:21) What’s next: monopsony and industrial policy * (1:17:59) Favorite philosopher: Rawls vs. John Stuart Mill * (1:19:45) Sign-off Justified Posteriors is the podcast that updates its beliefs about the economics of AI and technology, hosted by Andrey Fradkin and Seth Benzell. If we changed your priors, subscribe, share it with a friend, and keep your posteriors justified. Intro & Sponsor [00:00 – 00:47] [00:00:06] Seth: Welcome to Justified Posteriors, the podcast that updates beliefs about the economics of AI and technology. I’m Seth Benzell, excited to learn about what AI is other than what my bubbe says after I spill hot water on her, coming to you from Chapman University in sunny Southern California. Andrey: And I’m Andrey Fradkin, coming to you from San Francisco, California. We’re very thankful to our sponsors at Revelio Labs, purveyors of fine data products. And we’re very excited to have Ioana Marinescu join us today. Welcome to the show, Ioana. Ioana: Thank you. I’m so glad to be here. Make Policy Now: A Flexible Safety Net [00:47 – 05:51] [00:00:47] Andrey: To get started — you have this very provocative, interesting piece in the Digitalist Papers about various social policy solutions for transformative AI scenarios. Could you tell us about the piece? Ioana: Absolutely. As part of doing this Digitalist piece, I was thinking, as somebody who has worked a lot on the social safety net: what do we do if AI leads to a lot of job loss, like many people are saying it would? We’ll talk later about the various scenarios, but assuming that’s at least a possibility we have to acknowledge, what would you want to have from a policy perspective? And so I was really thinking hard about devising a flexible policy toolkit that will be able to address issues in the labor market no matter how big the shock is. That was the overarching theme of the policy design I’m proposing — just to start a discussion. I’ve tried to propose some helpful options, but it’s really with the idea of, let’s talk about doing something like this, what are the pros and cons. [00:02:10] Andrey: So what are the options on the menu for — let’s say AI comes along, a lot of people lose their jobs. The first thing we should get started with: do you think we should be making policy today, or should we wait until something happens and then make policy? Ioana: I think it’s very important to make policy today, but in a flexible way — meaning the policy cannot depend on some very specific detail of exactly how AI is going to impact the labor market, because we don’t know exactly what’s going to happen. It’s important to put the policy in place today because the political process is very long, so it may not be able to come online quickly enough when we really need it. That’s one reason. The other is — and I work a lot on social insurance — for workers, they want to and should feel insured. “Whatever happens, we the government have got you

    Ioana Marinescu on Insuring Workers for AI, Monopsony, and Philosophy
  6. 1. Juni

    Kevin Bryan on Bottlenecks, AI in China, and What Economists Should Actually Be Working On

    This week we to with Kevin Bryan, Associate Professor of Strategy at the University of Toronto’s Rotman School, author of the legendary economics blog A Fine Theorem, co-founder of the ed-tech startup All Day TA, and the man behind one of the most-discussed Twitter/X feeds in econ, @Afinetheorem. Kevin recently published a multi-book review of the economics of AI in the Journal of Economic Literature, and that’s where we start. Along the way we get into the gap between AI’s technical capability and its actual diffusion, the stages of how organizations adopt new technology, why the binding constraint on AI value is organizational integration (not prediction vs. judgment), what an AI-for-science research agenda should look like, the coffee test and the fence-post test, what forecasting surveys reveal about how economists and lab researchers actually differ, a dispatch from Kevin’s recent trip to China (spoiler: they are not AGI-pilled), the future of the academic paper, and a lightning round on comparative advantage in the age of AI. A wide-ranging, opinionated, very fun conversation. Grab your Chinese peptides and settle in. Links & References Kevin’s work * Kevin Bryan, “The Economic Impacts of Artificial Intelligence: A Multidisciplinary, Multi-book Review” — Journal of Economic Literature, 64(1), 2026. * A Fine Theorem — Kevin’s research blog * All Day TA — turn course content into a custom AI teaching assistant * Creative Destruction Lab — the accelerator Kevin helps run (first AI accelerator in the world, 2016) Books & essays discussed * Leopold Aschenbrenner, Situational Awareness — the essay Kevin gives all his students (”read chapter one, believe chapter one”) * Erik Brynjolfsson & Andrew McAfee, The Second Machine Age * Ajay Agrawal, Joshua Gans & Avi Goldfarb, Prediction Machines and the follow-up Power and Prediction * Joel Mokyr, The Gifts of Athena and A Culture of Growth — Kevin’s PhD advisor, “the Michael Jordan of progress world” People & projects mentioned * The Unjournal and Works in Progress — models for the “new journal” * Chad Jones, Stanford GSB — growth theorist read seriously by people in industry * Phil Trammell, GPI / Oxford — “Phil World,” the rapid-growth scenario * The coffee test (attributed to Steve Wozniak) and Kevin’s own fence-post test as benchmarks for embodied AGI Previously on Justified Posteriors * Avi Goldfarb — Prediction Machines, O-Ring Tasks, and How AI is Reshaping Economics * Alex Imas — Demand Collapse, Bargaining with Machines, and Behavioral AI Economics Our sponsor * This episode is brought to you by Revelio Labs, the leading provider of labor-economics data, available to academics on WRDS. Chapters * (00:00) Intro & sponsor * (00:39) The JEL book review: what the economics-of-AI canon got right — and what the older books still beat the new ones on * (03:19) Prediction vs. judgment, and the real bottleneck: organizational integration * (05:52) Too pessimistic on the tech, too optimistic on diffusion — Waymo, Pearl Street, and the COVID vaccine * (12:34) The four stages of how organizations actually adopt a new technology * (15:42) Status-quo bias, banning Anthropic, and treating frontier AI like nuclear material * (20:16) Why Situational Awareness beat the economists, and the book Kevin actually wants: AI for science * (26:53) Forecasting AI: the surveys, and where economists and lab researchers do (and don’t) diverge * (28:20) Benchmarks, the coffee test, and the fence-post test * (35:53) Rapid-growth scenarios, labor-force participation, and “Phil World” * (41:40) Scaling regularities: what economists should defer to technologists on — and what they shouldn’t * (43:34) Why forecasts matter for policy and capital allocation * (45:50) Dispatch from China: not AGI-pilled, “involution,” broken capital markets, EVs and self-driving * (1:01:40) War, nationalization, the end of open source — and why everyone in China uses Claude * (1:06:06) A Fine Theorem, the economics of blogging, and the rising value of taste * (1:17:48) The economist as plumber: comparative advantage, RCTs, and what grad students should do * (1:24:07) What the academic paper looks like in two years * (1:28:22) San Francisco, ambition, and the permission structure for growth * (1:32:56) Lightning round: favorite economists, All Day TA, and advice for econ grad students Open & Intro [00:00 - 00:39] [00:00:12] Seth: Welcome to the Justified Posteriors Podcast, the podcast that updates beliefs about the economics of AI and technology. I’m Seth Benzell, finally able to meet one of my theoretical heroes, coming to you from Chapman University in sunny Southern California. Andrey: And I’m Andrey Fradkin, coming to you from San Francisco. Excited to have Kevin Bryan as our guest today. Kevin, welcome. Kevin: Thanks for having me. Very excited. Andrey: Kevin is a leading thinker in the field of progress, and in AI economics. He also has his own startup, All Day TA, and is prolific on Twitter — at times. Kevin: At times. The JEL Book Review: What the AI-Econ Canon Got Right [00:39 - 03:19] Andrey: Kevin, you wrote an article reviewing several prominent books on AI. Why did you do this, and what did you learn from the exercise? [00:01:13] Kevin: It’s pretty interesting. Economics of AI is not that new of a field — some of the canonical books on how economics thinks about AI go back to before large language models existed. Books like The Second Machine Age by Brynjolfsson and McAfee, and Prediction Machines by Agrawal, Gans, and Goldfarb. These are pre-LLM — written before the attention paper. So it’s interesting to look at what of the core ideas in the economics of AI have changed given the technological improvements. On the technology side, I don’t think there have been massive surprises for people who were paying attention. At least since the scaling law paper, if you’d drawn the line on the graph, you’d have more or less predicted everything that happened. I remember reading Kurzweil — The Age of Intelligent Machines, The Age of Spiritual Machines — back in college, and those are just drawing different lines on the graph, in that case based on compute, and we’re getting very close to what actually happened. Likewise on the economic side: given that the technological trajectory hasn’t changed much, I don’t think the underlying economics has changed as much as people might think. Where things might be bottlenecked, how technology improvements map into growth, the effects on labor markets — the fundamental microeconomics of AI’s predictions hold up pretty well. I found it interesting how few of the 2023, 2024, 2025 books had really advanced my understanding of the economics of AI compared to the older ones. Prediction vs. Judgment, and the Real Bottleneck [03:19 - 05:52] [00:03:19] Seth: Lots to unpack. We just had Avi Goldfarb on the podcast and pressed him on his Prediction Machines approach, where he distinguishes the AI that’s good at predicting from the human that’s good at judging. If any of these books would have changed after gen AI, it’d be that one. Don’t you think that book maybe gets something wrong? Kevin: I think they’d agree — they wrote a follow-up in Power and Prediction. But the disagreement isn’t about the prediction-versus-judgment distinction. Even in the original book — and I remember talking to them about this in 2016, 2017 — judgment is a sliding scale. Take the umbrella example: I know my utility function on an umbrella, I know how much I dislike rain. I give the AI data, it looks at my face, sees light rain, heavy rain, and it can predict my utility function — in which case judgment is taken over by AI. Everyone understands that. That said, on the scale of how easy it is to figure out the underlying utility function from data versus the predictions that go into it, I don’t think that’s changed. None of the major language models technologically can — or even attempt to — modify how they operate for me versus you. They store a little memory and RAG their way into remembering what you’re like, but there’s no attempt to fine-tune the model. We’d like to use continual learning, but we can’t yet. So the judgment aspect is still pretty binding even today. Where I think there’s a difference — and where Ajay, Avi, and Josh would say they were wrong — is that the fundamental problem for AI’s creation of value isn’t prediction versus judgment. It’s the organizational integration problem. There’s overlap between the two, but we’d take the organizational and architectural bottlenecks more seriously now, partly because we’re applying AI to more complex tasks where those bottlenecks start to bite. Too Pessimistic on Tech, Too Optimistic on Diffusion [05:52 - 12:34] [00:05:52] Seth: You point this out with The Second Machine Age — Andy and Eric’s world-historical automated car ride. Andrey: It’s weird to think that in some ways they’re a little too pessimistic about the technology, but a little too optimistic about social diffusion. The driverless cars going down the highway in California are a perfect example. Kevin: Such a good example. We all talk to different audiences. When I talk to policy people, I tell them: “Whatever you think the capabilities of AI will be in the future — more than that.” This isn’t a sales pitch. Every single person inside the lab agrees. You have people high up in government who think about AI as the AI of today plus epsilon. And you want to ask: what did you see in the past 10 years that makes you think this is a good way to plan for the future? [00:07:01] On the other hand, out in California they wildly underrate diffusion friction. I give the Waymo example: if diffusion is so easy, how come we rode in a Waymo 10 years ago? I’m in Toronto — Jeff Hinton’s city — and there’s n

    Kevin Bryan on Bottlenecks, AI in China, and What Economists Should Actually Be Working On
  7. 19. Mai

    Seb Krier on AGI, the Coasean Singularity, and EDM

    Seb Krier on AGI, Scaffolding, and Coasean Bargaining at Scale In this episode of Justified Posteriors, we welcome Seb Krier — policy lead for AGI at Google DeepMind and excellent Twitter poster. Speaking in his personal capacity, Seb walks us through his understanding of AGI, why AI alignment has gone better than expected, the potential and limitations of a world where agents constantly barter on our behalf, and — of course — electronic music. We also cover AI in London vs. New York, how Seb went from reading Marginal Revolution for 15 years to becoming a recurring character on it, and Seb’s side-splitting humor on mediocre AI conferences. Related Links * Seb Krier on X: @sebkrier * Seb’s Substack, Technologik * “Coasean Bargaining at Scale” — Seb’s essay at the Cosmos Institute (also republished here) * “Musings on Recursive Self-Improvement” — Seb’s essay separating model-side RSI from societal-side * “The Cyborg Era: What AI Means for Jobs” — Seb’s guest essay on Alex Imas’s Substack, defending the scaffolding view * Anthropic’s Project Deal — the agent-bargaining experiment among Anthropic employees * Fradkin & Krishnan, “MarketBench” — Andrey and Rohit experiment of LLMs bidding in procurement auctions as an investigation of the future of AI marketplaces and the companion writeup: Rohit Krishnan, “Agent, Know Thyself! (and bid accordingly)” * Edge Esmeralda — Devon Zuegel’s pop-up village in Healdsburg, CA * MATS — for junior economists looking to skill up on AI safety/governance * Cosmos Institute and FIRE * bianjie.systems — the art platform Seb is co-organizing a dinner with in NY (Seb’s announcement) * Drexciya — James Stinson, Gerald Donald, and the Detroit electro-afrofuturism canon Timestamps (00:00) Intro (01:16) What is AGI? (07:30) In defense of scaffolding — Hayek, division of labor, and why one giant model won’t do it (13:00) Markets for cognition: will agents bid in procurement auctions? (18:40) Recursive self-improvement — separating the model side from the societal side (24:44) Alignment has gone better than 2017-Seb expected; prefer “intent following” (31:14) What economists should actually work on to inform AI labs(33:32) What does a DeepMind policy lead’s day look like? (38:20) AI Conferences(41:52) Coasean bargaining at scale — the positive vision(55:00) Inequality, property rights, and who gets the initial allocation (01:03:00) The Helldivers 2 “Managed Democracy” dystopia as Coasean bargaining gone wrong (01:09:00) Sponsor: Revelio Labs (01:09:30) Lightning round Justified Posteriors is a reader-supported publication. To receive new posts and support our work, consider becoming a free or paid subscriber. You’re also invited to our discord community at: https://discord.gg/b8VpPbBUt Transcript 00:00:00,100 --> 00:00:20,480 [Seth] [upbeat music] Welcome to the Justified Posterior’s podcast, the podcast that updates beliefs about the economics of AI and technology. I’m Seth Benzell, the number two biggest fan, after Tyler Cowen, in the Seb Krier fan club. 00:00:20,480 --> 00:00:20,740 [Andrey] [laughs] 00:00:20,740 --> 00:00:24,660 [Seth] Coming to you from Chapman University in sunny southern California. 00:00:24,660 --> 00:00:34,120 [Andrey] And I’m Andrey Fradkin, coming to you from San Francisco, California. And Justified Posterior’s is sponsored by the fine folks at Revelio Labs. 00:00:35,560 --> 00:00:45,600 [Andrey] We’re very excited to have Seb Krier here with us today. He is the policy lead for AGI at Google DeepMind, and is, 00:00:46,840 --> 00:00:52,400 [Andrey] dare I say, a thought leader in this space. Welcome to the show, Seb. 00:00:52,400 --> 00:00:54,200 [Seb Krier] Thank you very much. It’s great to be here. 00:00:55,380 --> 00:00:58,160 [Seb Krier] Yeah, I’m Seb, calling in from New York. 00:00:58,160 --> 00:01:00,320 [Andrey] And we should remind our listeners that 00:01:01,340 --> 00:01:08,410 [Andrey] Seb is, during this podcast, expressing his personal opinions, and is not speaking on behalf of DeepMind. All right. 00:01:08,410 --> 00:01:09,740 [Seb Krier] Indeed. [laughs] 00:01:09,740 --> 00:01:11,060 [Andrey] [laughs] 00:01:12,780 --> 00:01:13,900 [Andrey] The usual caveat. 00:01:15,260 --> 00:01:16,760 [Andrey] Seb, what is AGI? 00:01:18,080 --> 00:01:19,450 [Seb Krier] What is AGI? [laughs] 00:01:19,450 --> 00:01:19,570 [Andrey] [laughs] 00:01:19,570 --> 00:01:19,580 [Seth] [laughs] 00:01:19,580 --> 00:01:19,780 [Seb Krier] Great question. 00:01:19,780 --> 00:01:21,900 [Andrey] We’re going to start with the big questions. 00:01:21,900 --> 00:01:22,880 [Seb Krier] Yeah, might as well. 00:01:24,259 --> 00:01:54,840 [Seb Krier] [sighs] I think there’s so many definitions out there of what AGI is, and I think most of them are kind of unsatisfactory in one way or another. I’ve seen stuff like many definitions are indexed on the societal transformations or economic impacts of the technology, which I don’t really like very much because it makes it very dependent on external factors whether or not we have AGI. If it’s banned, we don’t have AGI, and if it’s not banned, we have AGI. Is it? 00:01:54,840 --> 00:01:55,480 [Andrey] [laughs] 00:01:55,480 --> 00:02:04,670 [Seb Krier] And there are other tests, like if an AI makes $1 million or something, which I find is very weird because most humans do not make $1 million in the first place. 00:02:04,670 --> 00:02:05,080 [Andrey] [laughs] 00:02:05,080 --> 00:02:11,359 [Seb Krier] So the one I kind of like is actually Shane Legg’s definition- 00:02:11,360 --> 00:02:11,620 [Andrey] Mm 00:02:11,620 --> 00:02:12,420 [Seb Krier] ... who’s at Deep Mind, who is 00:02:13,640 --> 00:02:16,980 [Seb Krier] more of a capability-based definition, which is something along the lines of 00:02:18,420 --> 00:02:20,960 [Seb Krier] an AI or a system that does most 00:02:22,380 --> 00:02:30,360 [Seb Krier] standard cognitive tasks that people typically do. [lips smack] So it’s kind of the bar isn’t too low, and it’s also not too high either. 00:02:32,220 --> 00:02:35,480 [Seb Krier] And so I think he’s got this definition of a minimal AGI, 00:02:36,580 --> 00:02:43,020 [Seb Krier] and I think that we’re not exactly there yet. I would disagree with people saying that we have AGI today because I think 00:02:44,220 --> 00:02:48,900 [Seb Krier] a lot of the systems we have, there’s many things that a human can do that they don’t really do very well. 00:02:48,900 --> 00:02:50,360 [Seth] What’s the biggest gap that we’re missing? 00:02:52,020 --> 00:03:47,740 [Seb Krier] I’d say there’s a few. One of them might be continual learning, or at least the ability to adapt and learn over time, and in different contexts and situations, just kind of update your own world model or whatever. If I think of a new joiner in a company, they’re not super useful the first day, but their value goes up over time because they learn all sorts of things. And so [lips smack] that might be one of them. A lot of the systems we have today, I think, are not very good at software, and you’re using graphical user interfaces and software and whatnot. If I ask an agent right now to go and use a music production software and make a track, I think they’d generally struggle. That doesn’t mean it’s impossible to solve or anything like that, but I think, in many respects, they’re not as general as you’d want them to be. And then the other bit also is, [lips smack] and of course they still make some silly mistakes here and there, but I think that’s getting it fixed. But the creativity point is one that I’m really interested in as well, in that I think they’re really good at kind of 00:03:48,780 --> 00:04:02,700 [Seb Krier] exploiting maybe an existing paradigm or an existing knowledge and so on, and recombining knowledge and whatnot. But I think really coming up with new concepts and abstractions entirely is something I think humans can do, but I don’t see our current systems really doing either. 00:04:02,700 --> 00:04:10,060 [Andrey] How do you measure whether humans can do creative tasks? One of the things that 00:04:11,200 --> 00:04:15,940 [Andrey] strikes me as a bit of an unfair test in that, 00:04:17,060 --> 00:04:23,290 [Andrey] let’s say you ask an LLM to write a poem or to write a story. It’s very- 00:04:23,290 --> 00:04:23,290 [Seth] [laughs] 00:04:23,290 --> 00:04:32,050 [Andrey] ... times more entertaining than what a random human would write. So, do you have a benchmark for creativity? 00:04:32,050 --> 00:04:35,390 [Seth] This is the meme where the robot asks Will Smith if he can compose an opera. 00:04:35,390 --> 00:05:14,700 [Seb Krier] [laughs] Can you? Yeah, exactly. It depends, and you’re right. Obviously, most people aren’t creating new abstraction and concepts on a day-to-day level. But I imagine there’s still something qualitative about that kind of creativity that I think does get applied in everyone’s day-to-day life in various kind of ways. Maybe they’re not as big or significant as creating a symphony. But I don’t really have a strong test. There’s actually an interesting podcast that had Ben Goertzel and Yoshua, I think a few years ago, where they were saying something like, if you had a model that was trained knowing only classical music and West African drumming, could it come up with jazz in the first place, or recreate jazz? 00:05:16,460 --> 00:05:27,880 [Seb Krier] And I quite like that test. And in principle, I can imagine it being possible. You could kind of decompose all sorts of different kind of elements and variables here and just get something jazz-like. But it still feels a bit... 00:05:29,580 --> 00:05:40,580 [Seb Krier] It’s not the same as just coming up with the idea of jazz in the first place and saying, oh, I’m goi

    Seb Krier on AGI, the Coasean Singularity, and EDM
  8. 4. Mai

    Avi Goldfarb on Prediction Machines, O-Ring Tasks, and How AI is Reshaping Economics

    This week, we’re joined by Avi Goldfarb, one of the leading economists of artificial intelligence and co-author of Prediction Machines. Avi has been thinking seriously about AI economics long before the ChatGPT shock, so we asked him what he thinks the earlier framework got right, what it missed, and how economists should update their beliefs now. The conversation starts with Avi’s seminal book, Prediction Machines, and the idea that AI is best understood as a drop in the cost of prediction, which is a complement to judgement. We ask what that book got right and what it got wrong. From there, we interrogate Avi on the murky boundary between prediction and judgment. We had investigated the idea that maybe judgment and prediction were not as separable as economists like to believe in our episode with Alex Imas. We also ask whether, if AI gets better at predicting human judgment, whether judgment disappears, or do humans simply “move up the stack”? And what is taste exactly? Avi says that sometimes judgment becomes predictable, but humans still matter because goals, values, organizational politics, and “what matters” are often implicit, unstable, and hard to codify. Avi shoots down Seth’s galaxy-brain suggestion that correct ontology choice — i.e., deciding what sort of natural kind a thing is, or understanding when a problem is out of context — is a uniquely separate skill (taste?), calling it just another prediction error. But he does concede that deciding how much to prepare for ‘Black Swan’ events may be an enduring role for judgment. We then revisit the O-ring theory of production and what it means for automation. We had covered Kremer’s article in a recent episode (see here) and asked Avi about his new paper, riffing on the idea at the worker level. Avi says that if tasks inside jobs are complements rather than substitutes, then automating one task may make the remaining human tasks more valuable, not less. Avi explains why workers may reallocate attention toward the tasks machines cannot yet perform (shooting down Seth’s suggestion that this is actually difficult in most jobs). The discussion also covers whether AI will augment or replace workers, whether governments should try to steer AI toward human-complementing technologies, and why that distinction may be much harder to define in practice than it sounds. Avi agrees with Andrey and Seth’s pushback on “augmentation good, automation bad” framings (e.g. friend of the show Erik Brynjolfsson’s “Turing Trap”). Then we get into forecasts: how fast AI capabilities might advance by 2030, what that means for GDP growth by 2050, whether GDP is still the right thing to forecast, and why even very powerful AI may run into bottlenecks in the real economy. We use the paper Forecasting the Economic Effects of AI to ground the discussion. We close with lightning-round topics including AI’s impact on centralization, privacy/de-anonymization, peer review, and whether academic journals still serve the function they once did. Papers, books, and ideas mentioned * Avi Goldfarb’s seminal book with Ajay Agrawal, and Joshua Gans — Prediction Machines * A black swan is the occurrence of a wildly unpredictable event, which Nassim Taleb argues, in his book by the same name, is more common than we like to think * A New Riddle of Induction — by Nelson Goodman — is the source of Seth’s thought experiment about “bleen”, a color which is green until 2029 and blue after, and green * Michael Kremer — “The O-Ring Theory of Economic Development”, covered in this episode of the pod: * Daron Acemoglu and Pascual Restrepo’s task-based models of automation, especially “The Race Between Man and Machine.” * Avi mentions David Autor and Ben Thompson on automation and skill scarcity when Seth comments that you may not be able to reallocate effort between tasks as a worker, including their paper “Expertise” * Erik Brynjolfsson in the “Turing Trap” argues that automation technologies are less good than augmenting technology * Eric Topol’s book on AI in medicine — Deep Medicine * John Markoff — Machines of Loving Grace — The source of a title for an influential essay of the same name by Dario of Anthropic. Both draw from an earlier poem about a Sci Fi utopia: https://allpoetry.com/All-Watched-Over-By-Machines-Of-Loving-Grace * Korinek and Stiglitz on AI, capital, and taxation; Lockwood and Korinek on optimal taxation and automation — We covered these topics at the end of our episode with Basil Halperin in the context of “Tax Policy at the End of History” around the 1:19:00 mark * We talk about de-anonymization, and Avi references this provocative paper from Florian Ederer * Avi brings up Bob Gordon, and his argument, famously in the book The Rise and Fall of American Growth, that the early 20th century was incredibly important for increases in US living standards, which digital technologies have not lived up to * Digital Hermits, by Jeanine Miklós-Thal, Avi Goldfarb, Avery M. Haviv & Catherine Tucker, is a paper by Avi thinking about how information spillovers, now from AI, drive some people to be more private than they would otherwise be. In our conversation, we speculate AI will make these hermits even more “hermetic” * We discuss this paper on new forecasts of AI and its impact on economic growth: Forecasting the Economic Effects of AI * Refine and AI-assisted peer review are discussed in this pod. For more, see our episode with Ben Golub, founder of Refine. This episode is sponsored by Revelio Labs — a great source of labor economics data for academics and firms. Now available on WRDS. Join our Discord community at this link: https://discord.gg/w3GSapx2d Transcript Introduction [00:00] Seth: Welcome to the Justified Posteriors podcast, the podcast that updates beliefs about the economics of AI and technology. I’m Seth Benzell, your loyal non-fiction machine, coming to you from Chapman University in sunny Southern California. Andrey: And I’m Andrey Fradkin, coming to you from San Francisco, California. And we are very happy that Justified Posteriors is sponsored by the fine folks at Revelio Labs. And we’re very delighted to have Avi Goldfarb, who is a leading thinker in the field of AI economics and has also been a personal mentor on the show. We’re very excited to hear his thoughts on a variety of topics. Welcome, Avi. Avi: Thanks so much and thanks for having me on the show and looking forward to it. Andrey: All right, let’s get started. I have in front of me this book that you might remember writing at some point. Seth: Gaze into the soul of the man in the bookstore. What Did Prediction Machines Get Wrong? [01:12] Andrey: Now, I just think it’s a good cover. And I had to check: when was it released? It was released in 2018. And as I was skimming through it, you know, a lot of interesting points made there are still things that we’re talking about today, almost 10 years after it was released. So let me start off with the following question. And then maybe we can work backwards more into the ideas in the book. But what do you think prediction machines got wrong? Avi: I think prediction may... I’ll start with a hard question. Seth: No softballs on Justified Posteriors. Avi: So on the specifics of which industries and when, to the extent we tried, at least I did not anticipate how quickly language and coding would become prediction problems. And when we talk about disruption and industry disruption, a lot of the examples are things like driving, and we talk about radiology. And we still have plenty of radiologists around. Self-driving cars and trucks. seem like they’re now imminent, but it certainly took a lot longer than we expected back in 2018. Andrey: So is it a fair assessment to say that the large language models, even in 2018, weren’t on your radar? I guess they weren’t on many people’s radar. The Three Ideas of Prediction Machines [02:45] Avi: Not really. We have some discussion of machine translation. So that’s in there as a huge potential use case, but the arrival of ChatGPT and how it sort of changed how we interact with machines and how we think about AI was not really there. Another way to put it is prediction machines had three ideas. So idea number one is AI can be framed as a drop in the cost of prediction. So prediction. As in filling in missing information, statistical prediction is getting better, faster and cheaper. Idea number two is that when something gets cheap, you start using it for unanticipated uses. So when arithmetic got cheap, it wasn’t just that we use computers for accounting. We started to use computers for all sorts of things that we never used to think of as arithmetic problems like imaging and mail and music. And then idea number three is what are the complements to machine prediction? And we talked about data and judgment. The book, and certainly our attention to the book in the first three or four years after it was published, was on idea number one and idea number three. So identify prediction problems in your organization, and then think about what data you need to make those predictions better, and try to understand what matters to you in terms of judgment. And that second point kind of got lost. But in the last four years, it’s become clear to me is that that second point was maybe the biggest one, which is this tool, which still under the hood is computational statistics, enables us to find all sorts of applications for computational stats that we didn’t really imagine before. Judgment and data are still gonna be useful, but that phase one, that step one, that first idea of identifying prediction problems, that’s not really how we think about using AI today. And in some sense, that... was a missing emphasis throughout the book and throughout how we thought about that book, or at least how I thought about that bo

    Avi Goldfarb on Prediction Machines, O-Ring Tasks, and How AI is Reshaping Economics

Info

Explorations into the economics of AI and innovation. Seth Benzell and Andrey Fradkin discuss academic papers and essays at the intersection of economics and technology. empiricrafting.substack.com

Das gefällt dir vielleicht auch