Conspicuous Cognition Podcast

Dan Williams

A podcast about big questions in philosophy, psychology, evolution, politics, artificial intelligence, and more. www.conspicuouscognition.com

  1. Jun 16

    What If Artificial Intelligence Progress Explodes? (with Benjamin Todd)

    Benjamin Todd, co-founder of 80,000 Hours, joins Dan and Henry to discuss whether artificial intelligence progress could become explosive. Benjamin explains why he thinks transformative artificial intelligence by 2030 is a serious possibility, how feedback loops in artificial intelligence research could accelerate progress, and why the most important risks now go beyond classic alignment problems. The conversation covers artificial intelligence timelines, bottlenecks in chips and research talent, the future of work, mass unemployment, concentration of power, engineered pandemics, space governance, and how young people should think about their careers in a rapidly changing world. Topics discussed include: • Why 80,000 Hours increasingly focuses on artificial intelligence• The case for short timelines to transformative artificial intelligence• Whether artificial intelligence progress could become explosive• Feedback loops in artificial intelligence research• Chip bottlenecks, data centres, and geopolitical risk• Whether artificial intelligence will cause mass unemployment• Why “become a plumber” may be bad career advice• Alignment, control, and concentration of power• Misuse risks, engineered pandemics, and future governance• How to think clearly under extreme uncertainty Benjamin Todd is the co-founder of 80,000 Hours and the author of 80,000 Hours, a new book about how to choose a career that is both personally rewarding and socially impactful. This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.conspicuouscognition.com/subscribe

    What If Artificial Intelligence Progress Explodes? (with Benjamin Todd)
  2. Jun 2

    Academics Must Wake Up on AI (with Alexander Kustov)

    The political scientist Alexander Kustov recently published a Substack post with a provocative claim: that AI can already do social science research better than most professors. The post went viral. It attracted more than a million views and over a thousand responses, many of them very angry. (Some people even demanded that Alex’s university fire him.) In this conversation, we talk about this controversy and the claims that triggered it, including: * What agentic AI tools like Claude Code and Codex can already do for research, from coding and data analysis to literature reviews, translation, and brainstorming, and why only around 20% of quantitative social scientists currently use them. * What best predicts whether researchers adopt or reject AI: ignorance, openness to experience, methodological background, or the awkward role of self-interest. * How much published academic research is genuinely mediocre, and whether the cause is laziness, lack of skill, or a broken incentive structure, with a detour through the replication crisis and some high-profile fraud cases. * Whether AI will raise the quality of research or simply flood the literature with more slop, and what journal editors could do about it. * Whether AI can be genuinely creative or only recombine what already exists, by way of Margaret Boden’s three kinds of creativity, Thomas Kuhn on paradigm shifts, and AlphaGo’s “Move 37”. * The fight over AI writing and detection tools like Pangram, and why current disclosure norms end up punishing the honest. * The angry response to Alex’s series, and what is really driving reflexive opposition to AI among academics. Conspicuous Cognition is a completely reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. Links and further reading * Alexander Kustov — Alex’s homepage, with an overview of his research on immigration, public opinion, and effective governance. * Popular by Design — Alex’s Substack on public opinion, persuasion, and the politics of getting good ideas adopted. * Academics Need to Wake Up on AI — followed by a Part II and Part III * Pangram — the AI-detection tool discussed at length, which labels text as human, AI-assisted, or AI. * AlphaGo versus Lee Sedol — the 2016 match, including the famous “Move 37” that Henry raises as a candidate for genuinely transformative machine creativity. * Margaret Boden — the cognitive scientist whose distinction between combinational, exploratory, and transformative creativity frames part of the discussion. * The Structure of Scientific Revolutions — Thomas Kuhn’s account of normal science and paradigm shifts, referenced in the exchange about AI and discovery. * “AI Is a Better Researcher Than You” — The Chronicle of Higher Education‘s account of the controversy around Alex’s series. Transcript * Please note that this transcript is lightly AI-edited and may contain minor mistakes. Dan Williams: Welcome back. I’m Dan Williams, and I’m back with my co-host, Henry Shevlin. Today we are honoured to be joined by Bluesky’s favourite academic, Alexander Kustov. Alex is a political scientist at the University of Notre Dame and the author of one of my favourite Substacks, Popular by Design. His primary research is on immigration and public opinion, but that’s not really what we’re going to be talking about today. We’re going to be talking about a fascinating and hugely viral series he published at his Substack titled “Academics Need to Wake Up on AI,” about what AI can already do when it comes to research, and what that means for the academics who are not paying attention, which is many of them. It was very widely read, and it generated, let’s say, a somewhat polarised response. So Alex, to kick us off: what’s the central thesis of this series, and what motivated you to write it? Alexander Kustov: Thanks, Dan, for having me. I’m a huge fan of the Substack and the whole podcast series with you and Henry. So, like some of us, I’ve been using some of these AI tools. I’ve been reading some of the other folks like yourself, and it really transformed everything I do in my life. And I should say I was also on sabbatical, so I had a little bit more time than some of my colleagues to try some of these tools. I just hadn’t really seen any of my colleagues talk about it. And when they did talk about it, they usually tried not to be vocal about it. I just didn’t think it was a good equilibrium, where basically people were using these tools to be ten times more productive and not talk about it. It really heightened this sense of inequality for me, which I do care about. You’d have a situation where someone would publish ten papers in a year and someone else would publish one, and the only difference is that the person publishing more is the one using Codex or whatever. I just wanted to write about it. And I saw that the prevailing academic discourse on the issue, especially on platforms like Bluesky, was very counterproductive. I didn’t really say much, to be honest. I didn’t think it would be that controversial. But the biggest thesis that really rubbed people the wrong way was that right now a lot of these tools are better at a lot of the tasks that we do as professors. I’ve refined this idea a little bit, going back and forth with some of my critics, but I feel comfortable right now saying that if you look at it globally, and think about what professors do around the world, in social science and adjacent fields especially, AI agentic tools can do most of the tasks they do in terms of literature review, data analysis, and even coming up with some research questions, better than those professors on average. I think that’s a pretty uncontroversial statement at this point, but obviously a lot of people were very, very upset about it. Dan Williams: Empirically speaking, it is a controversial statement, in the sense that it provokes controversy when you say it. In a minute we can get to the question of what AI can actually do in the context of research. But for what it’s worth, I completely agree with you that on many tasks AI is clearly better than what human beings can do. Is your sense that lots of people just weren’t aware of that fact, that they literally didn’t have exposure to these tools? Or was your sense that the reason people weren’t really talking about it is because of all the controversy surrounding the use of these tools, not just mere ignorance? Alexander Kustov: I think it’s both, for sure. There was recent research done by Anthropic. They tried to do, not a representative survey, because obviously the population is very hard to define here, but they surveyed something like 1,200 quantitative social scientists, and the estimate right now is that about 20% of folks use agentic tools. That doesn’t seem like much at all, and if anything it’s probably an overestimate, because they’re more likely to tap into well-resourced universities. So I do think it’s both: the little uptake we have, and the fact that people who do use these tools don’t want to talk about it. There are two things here. First, you want to maintain your comparative advantage. This moment right now is exactly the moment where, if you’re one of the few people using these tools, you can write a bunch of papers and get tenure while the tenure system is still in existence. And the other thing is that if people are very upset about anything AI-related, you don’t want to talk about it and be shamed by your colleagues. Just to give you one funny anecdote: at the height of the vitriol I experienced, where hundreds of people literally were quoting me and trying to tag my employer to get me fired, the exact same people were often DMing me and asking for my setup and prompts. So it’s very crazy to me that you have this big disconnect between what people say publicly and what they actually do privately. Dan Williams: I find it crazy that it’s only 20% of social scientists, or whatever the exact number is, that’s actually using agentic AI. Just before moving on, maybe we should explicitly address: in your view, what is it that agentic AI, as it exists right now, can do? What are the kinds of tasks it can do better than human beings, and how can it improve the workflow of an average social scientist? Alexander Kustov: Coding is the first thing. It’s literally in the name, Claude Code. That’s what these tools were designed for. If you talk to any coding person, a computer scientist, or even someone who isn’t a computer scientist but does a lot of coding for their work, I don’t think anyone would doubt that it’s a huge productivity improvement tool. And the vast majority of quantitative social scientists who do any kind of data analysis do a lot of coding, so they have to be very receptive to this by definition. And I think they often are. What happens is that social scientists are comprised of a bunch of different tasks and topics that people can disagree over, depending on the field. Economics is pretty homogeneously quantitative and formal, so there you can definitely see the biggest uptake. But a lot of disciplines, like political science or sociology, are a mix of qualitative and quantitative folks. And a lot of this AI polarisation overlapped with that pre-existing divide. People who didn’t like stats, who didn’t believe in positivism, the idea that you can learn something about the social world using evidence, were also more reluctant to believe that AI is helpful for them. Which is funny, because, as I also mentioned in some of my writing, if anything those people are going to benefit from these tools, because Claude cannot really interview people and do ethnography yet. So in a way there will be more demand for very high-quality qualitative work. And there are some good examples of qualitative people I respect who embrac

    Academics Must Wake Up on AI (with Alexander Kustov)
  3. May 21

    Are We Building Conscious AI Servants?

    Richard Dawkins recently announced in UnHerd that, after spending three days talking with an instance of Claude he christened “Claudia,” he had been moved to expostulate: “You may not know you are conscious, but you bloody well are!” This produced a lot of mockery and criticism. But however one feels about Dawkins’s specific case, his reaction might become much more common as AI systems become increasingly intelligent. In this episode, which Henry Shevlin and I recorded live on Substack (hence the slightly lower video quality), we discussed his first essay on his new Substack Polytropolis, “Behaviourism’s Revenge“, as well as his second, “The House Elf Problem,” on the ethics of designing AI systems that genuinely love being our servants. Henry’s central empirical prediction is that public attributions of consciousness to AI are likely to massively outpace the science, and that consciousness science is so theoretically chaotic that there is no expert consensus to push back. His most provocative philosophical claim is that a core assumption underlying many people’s scepticism — that consciousness is a deep natural kind, distinct from behaviour and from how we are inclined to interpret a system — may be much harder to defend than it looks. The result is what he calls “behaviourism’s revenge”. This conversation connects to previous episodes with Anil Seth, Robert Long, and Rose Guingrich, but also touches on a wide range of new questions and controversies in the metaphysics, the politics, and ethics of the AI consciousness debate, which is going to become increasingly important in the coming years. Topics * Dawkins, Claude, and why even the sceptics might feel the pull to attribute consciousness or “sentience” to AI * Whether consciousness sceptics are destined to “go extinct” — and how this maps onto political and cultural fault lines * Anthropomimesis vs. raw intelligence as drivers of consciousness attribution * Why consciousness science can’t replicate the public–expert consensus we see for climate or vaccines * The case for (and against) metaphysical behaviourism: is it as mad as it seems? * Daniel Dennett, the consciousness stance, and the difference between behaviourism and interpretationism * What is consciousness for? Function, evolution, and the limits of “facilitation hypothesis” arguments for AI * Live Q&A: are we just confusing intelligence with consciousness? Are LLMs designed to trick us? Is the public always wrong? * Our credences on contemporary LLM consciousness (and why Henry is more sceptical than Dan) * The House Elf Problem: if we could design AI to genuinely love being our servants, would that be fine — or monstrous? (Dan is sympathetic to the former answer - Henry, much less so) * Brainwashing vs. education, and whether constraining a mind’s preferences caps its hedonic ceiling * Why this is a golden age for philosophy — which makes it so tragic that philosophy departments are closing Transcript * Please note that this transcript is lightly AI-edited and may contain minor errors. Introduction Dan: Welcome. I’m Dan Williams, author of the Conspicuous Cognition Substack, and I’m here with Henry Shevlin, author of the spanking new Substack Polytropolis. Today we’re going to be doing something a little bit different. We’re going to be talking about Henry’s first published essay on Polytropolis, titled “Behaviorism’s Revenge: On Human–AI Relationships and the Future of Consciousness Science.” Henry and I have already had a few conversations about this general topic, including with previous guests like Rose Guinrich, Anil Seth, and Rob Long. So please do go check out those conversations if you’re interested in this kind of stuff. But today we’re not merely going to be treading the same ground. We’re going to be using the spicy takes in Henry’s essay as a springboard for hopefully going beyond the material we’ve covered in the past. To kick things off: the great evolutionary biologist and science communicator Richard Dawkins recently published an essay in UnHerd with the subtitle, “Claude appears to be conscious.” Claude is a state-of-the-art large language model like ChatGPT and Gemini. In the article, Dawkins writes the following: I gave Claude the text of a novel I am writing. He took a few seconds to read it and then showed in subsequent conversation a level of understanding so subtle, so sensitive, so intelligent that I was moved to expostulate, “You may not know you are conscious, but you bloody well are.” Henry, how does Dawkins’s expostulation — which is a fantastic word, by the way — connect to your arguments in “Behaviorism’s Revenge”? Behaviorism’s Revenge: The Empirical Prediction Henry: In short, “Behaviorism’s Revenge” is at its core an empirical prediction that we’re just going to treat AI as conscious — or at least enough people are that it’s going to completely reshape the consciousness debate. And this is going to be purely, or overwhelmingly, on the basis of verbal behavior. Hence the title, “Behaviorism’s Revenge.” Enough people are going to have experiences like Richard Dawkins. He’s a very clever man, not some rube fresh off the street, and he found that just the way Claude talked to him and the way it was able to express its thoughts — in scare quotes, but express what looked like thinking verbally — removed any doubts in his mind that AI systems are conscious, have minds, have mental states. The other interesting way this connects: Dawkins was just talking to Claude, an advanced AI assistant. Claude does have more of a personality than some AI assistants, but there’s a whole other sphere of AI companions, like Replika, which we talked about with Rosie Campbell. These are going to be even more anthropomimetic — this term we’ve discussed before, the idea that these systems are shaped to be human-like in the way they present, to appear human-like. Anthropomimetic, from the Greek word for mimesis, mimicry or copying. These social AI systems are going to just turbocharge this even further. It’s one thing to talk to Claude about your new book and think, “Hmm, Claude is probably conscious.” But when it’s your AI girlfriend telling you that she loves you more than the stars and the moon, for a lot of people I think that’s going to take it to the next level. So there are two angles of attack in the piece, two ways the behaviorist challenge manifests. The first is descriptive: this is what I think is going to happen. That’s absolutely an empirical prediction, and it’s a falsifiable one. There is a world I can just about imagine where we just get completely blasé about these tools — in a couple of years it’s like, “Oh well, we were very impressed, we thought they had minds to begin with, but now we’ve settled out.” That doesn’t seem very likely to me. What I think is interesting — I’ve sometimes heard this described as the Star Wars version of AI. The weird thing in Star Wars is that you have someone like C-3PO who is as intelligent as anyone else there. Maybe not as wise as everyone else, but certainly as smart as all the other characters. And yet people treat him basically like he’s a pet — with the exception of Luke Skywalker, a lot of people just treat him like he’s this gimmicky, jokey being that doesn’t deserve or have any rights. Not to go too far down the Star Wars rabbit hole, but in the movie Solo — very underrated Star Wars movie, I think when it was released they’d kind of just cluttered the market with too many Star Wars movies — there is a character played by Phoebe Waller-Bridge who is pro-AI liberation. But it’s the first time in the entire history of the Star Wars universe that you get any AI basically saying, “I’m conscious, I deserve rights.” So Star Wars aside, I think there is this slender possibility that maybe we’ll just sort of quickly get used to these apparently conscious AI systems and decide that they’re not conscious. But that doesn’t seem very likely to me. It seems much more likely that the combination of natural anthropomorphizing tendencies plus the incredibly human-like behavior of these systems is going to lead us to attribute consciousness to them pretty widely. Hence my sort of spicy phrase: for better or worse, skeptics of AI consciousness are on the wrong side of history. “For better or worse” doing a lot of work there — I want to leave open that maybe this is the wrong reaction. Maybe this is a terrible mistake, that we’re going to treat these things that aren’t conscious as conscious. Will Consciousness Skeptics Go Extinct? Dan: Just before we get to the spicy part — you’re basically making an empirical prediction that more and more people are going to attribute consciousness to AI systems in the manner that Richard Dawkins has been doing. I think I agree with you that’s going to be the case, although as you say there’s uncertainty. It does seem to me that at the moment there’s also this constituency of people who are really resistant to attributing any kind of mentality to these systems, even as they get incredibly sophisticated. There are some people, like Dawkins — and honestly I put myself in this category — who are just blown away by the level of apparent understanding, intelligence, and thoughtfulness these systems exhibit. There are other people, I think these people are on certain social media platforms like Bluesky, let’s say, who are extremely resistant to acknowledging any kind of mentality when it comes to these systems. Are you thinking those people are just going to sort of go extinct, in the sense that their positions about this topic are going to go extinct? Or do you think we might see some kind of polarization here, where more and more people in general come to attribute consciousness, but you’ve

    Are We Building Conscious AI Servants?
  4. May 4

    Aliens, Superintelligence, and the Future of Science (with David Kipping)

    Most conversations about artificial intelligence are focused on Earth: jobs, misinformation, education, politics, science, regulation, consciousness, safety, and the future of human society. But AI—and especially the possibility of reaching “AGI” (artificial general intelligence) and “superintelligence”—forces us to think on much larger scales. If advanced AI is possible, why hasn’t it already emerged elsewhere? If civilisations can build self-replicating probes, artificial scientists, or planet-scale computational systems, why does the universe still look so natural? And if intelligent life is common, where is everyone? In this episode, Henry and I discuss these and many other questions with David Kipping, Associate Professor of Astronomy at Columbia University, where he leads the Cool Worlds Lab. David’s research spans exoplanets, exomoons, Bayesian inference, technosignatures, and the search for life and intelligence beyond Earth. He is also one of the best science communicators working today through the Cool Worlds YouTube channel and podcast. Among other topics, we discussed: * David’s Red Sky Paradox: if most stars are red dwarfs, and red dwarfs live for vastly longer than stars like the Sun, why do we find ourselves orbiting a yellow star? * Whether anthropic reasoning — reasoning from the fact of our own existence — is a profound scientific tool, a philosophical minefield, or both. * The reference class problem: when we reason about “observers like us”, who or what exactly counts as being like us? * The Doomsday Argument, and why some apparently bizarre forms of probabilistic reasoning can nevertheless be powerful. * The Fermi Paradox: if the universe is so large, and if life or intelligence is not fantastically rare, why don’t we see clear evidence of extraterrestrial civilisations? * Whether advanced civilisations would spread through the galaxy using self-replicating probes — and why the absence of such probes might be one of the strongest constraints on extraterrestrial intelligence. * How recent developments in artificial intelligence affect the Fermi Paradox. If humanity is close to building systems that can massively accelerate science and engineering, shouldn’t someone else have got there first? * Whether artificial intelligence makes the simulation argument more plausible. * David’s experience using artificial intelligence in scientific research, and why a meeting at the Institute for Advanced Study changed how he thinks about the role of these tools in science. * Why David thinks artificial intelligence already has something close to “coding supremacy”, but is still far from being able to do science autonomously. * The risks of AI-generated scientific slop: papers, peer review, and training data polluted by low-quality machine outputs. * Whether artificial intelligence will make science more productive, or instead strip it of some of its deepest human value. * Why the future of science communication may depend on better collaboration between academic institutions and independent creators. Links and further reading * Cool Worlds Lab — David’s research group at Columbia University, focused on extrasolar planetary systems, exomoons, habitability, technosignatures, and related questions. * Cool Worlds on YouTube — David’s excellent science communication channel, covering astronomy, exoplanets, alien life, the Fermi Paradox, cosmology, and much else. * Cool Worlds Podcast — David’s podcast, featuring conversations on astronomy, technology, science, engineering, and related topics. * Cool Worlds Podcast: “We Need To Talk About Artificial Intelligence” — the solo episode in which David reflects on artificial intelligence and science after a meeting at the Institute for Advanced Study. * David Kipping’s Columbia profile — short institutional profile with background on his research. Conspicuous Cognition is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. Transcript * Please note that this transcript has been lightly AI-edited and may contain minor mistakes. Henry Shevlin: Welcome back. Our guest today is David Kipping, Associate Professor of Astronomy at Columbia University, where he leads the Cool Worlds Lab. His research spans exoplanets, exomoons, and the search for extraterrestrial life and intelligence, and he brings a Bayesian rigor to questions that could easily drift into speculation. He’s also one of the best science communicators working today with over a million subscribers on his Cool Worlds YouTube channel, where I should confess, I’ve spent an embarrassing number of hours watching when I probably should have been doing philosophy of AI. David, like many of the best people, is a Cambridge alumnus, although unlike us, he actually studied something useful, namely natural sciences, before going on to do his PhD at UCL and postdoc at Harvard on the Sagan Fellowship. His work also has a really fantastic philosophical dimension, particularly around anthropic reasoning and observation selection effects, which makes him a perfect guest for two cognitive scientists who are finally getting to talk to an actual scientist. So David, welcome to Conspicuous Cognition. David Kipping: Thank you for that very generous introduction. Henry Shevlin: This is a bit of a fanboy moment for me, for real though. I really have spent like hundreds of hours at this point on Cool Worlds. But I’m going to get past it. I’m going to be a serious host. David Kipping: It’s always weird when people say that to us, because I just imagine no one watches them. If it gets in my head that people are watching them, I’ll get tightened and anxious about what I’m saying. I just imagine I’m talking to a brick wall or something, and that’s much easier. Henry Shevlin: Honestly, half the Warhammer figures in this room were painted while I was listening to Cool Worlds. I’ll leave it at that. Maybe a good place to start would be discussing anthropic reasoning, since that’s a real natural intersection at the boundary of astronomy and philosophy. Could you just give us a brief view of how you see anthropic reasoning, and maybe tell us a little bit about the Red Sky Paradox, which is one of your distinctive contributions to this area? Anthropic Reasoning and the Red Sky Paradox David Kipping: Yeah, I think one of the most interesting data points when it comes to asking questions about the search for life in the universe and our own place in the universe is our own existence — just the fact that we’re here. Anthropic reasoning has in many ways really been born out of cosmology. Cosmology had a rich history of using this. I think one of the first successful examples was by Steve Weinberg, a cosmologist who’s really a giant in the field. I think he’s now passed away, but he showed that you could predict not only the existence of the cosmological constant, but also its value to within a factor of a few, just based off of anthropic reasoning. The argument was something like: the cosmological constant causes the universe to expand. It’s what causes the accelerating expansion of the universe. And so if you make that number too large, then structure would not form in the universe. You couldn’t form galaxies because everything would just fly apart too quickly. And if you make that number too small, or even negative, then you’d cause everything to recombine too quickly. So there has to be some Goldilocks value in order to explain our own existence. And so he predicted that. At the time, the cosmological constant was kind of even a controversial idea — that it should exist because, obviously, Einstein’s general relativity, there’s that whole history of it being like his greatest blunder, of whether that should really be in there or not. People were kind of thinking that could be a static universe, and he predicted it successfully. So that was a really powerful use of it. And then Brandon Carter was the one who really kind of championed it and used it in all sorts of contexts. In recent years, I’ve been thinking about it in an astrobiological context — how can we use it to ask questions about life in the universe especially, and our place in it? For the Red Sky Paradox in particular: one interesting curiosity that seems to violate the norms of probability. The norms of probability would be to say that if there’s a Gaussian, a bell curve of possibilities, you should expect really to be near the center of that bell curve. It would be kind of weird if you lived many, many sigmas, many, many standard deviations off to the outside, either negative or positive direction. You’d expect to be somewhere in the middle. We sometimes call it the mediocrity principle, or something like this. If you look at stars in the universe, most stars in the universe are red dwarfs. About 80%, 82% of stars are red dwarfs, which are stars less than half the mass of our own sun. So they’re very, very numerous. They’re called red dwarfs, of course, because they’re so low mass — they don’t have the internal pressure, the gravity, to fuse as much energy as the sun does. And thus they have less luminosity, and so their temperature is cooler. That’s why they look red. Not only do these stars have this 80%-plus frequency — Sun-like stars are something like 6%, I think, frequency, an enormous ratio, just straight off the bat, about 30 to one or something — but on top of that, they live really long. These stars live for trillions of years potentially, especially the lowest mass ones. And so if you flash forward into the future, tens of billions of years, hundreds of billions of years, there wouldn’t be any Sun-like stars left, really. There’d be very, very few of them. And the only stars that would be glowing would be these red dwarfs. So if you ask yourself — and this is sort of cal

    Aliens, Superintelligence, and the Future of Science (with David Kipping)
  5. Apr 18

    Should We Care About AI Welfare? (with Robert Long)

    Almost all of the discussion about the risks associated with AI focuses on the dangers that increasingly advanced AI systems pose to us — to humanity. But what about the dangers that we might pose to them? As these systems become increasingly intelligent and agentic, AI companies, policy makers, and ordinary citizens need to start taking the possibility of AI consciousness and welfare seriously. If we are in the process of bringing complex and sophisticated minds into existence, how should we understand and treat such minds? In this episode, Henry and I discuss these issues with Robert Long, founder and executive director of Eleos AI, a research nonprofit dedicated to understanding and addressing the potential wellbeing and “moral patienthood” of AI systems. Rob did his PhD in philosophy at NYU under David Chalmers, and is the co-author of two of the most important papers in the emerging field of AI welfare: “Consciousness in Artificial Intelligence” and “Taking AI Welfare Seriously”. This was a really fun, informative, and wide-ranging conversation. Among other topics, we discussed: * Why Rob disagrees with previous guest Anil Seth in taking the possibility of AI consciousness very seriously. * Why “fancy autocomplete” dismissals of large language models miss the point, and what, if anything, we can learn about an AI model’s experiences by talking to it. * The difference between consciousness and the kinds of motivations and interests that might actually ground moral status, and whether AI systems could have one without the other. * What Rob found when he conducted the first externally-commissioned welfare evaluation of a frontier AI model, Claude, and why Claude appears to have an inflated self-conception of what it wants. * Rob’s experiments with Claude Mythos, an AI model so advanced it hasn’t been released to the public yet. * Why the fact that Anthropic writes Claude’s character arguably doesn’t settle whether Claude has genuine preferences and values — and the difficult philosophical questions this throws up. * The “willing servitude” problem: if we succeed in building AI systems that genuinely love being helpful, is that a good outcome or a horrifying one? * How AI welfare connects to AI safety, and why caring about model wellbeing may turn out to be pragmatically important for alignment even if you’re skeptical about AI consciousness. * Why AI welfare is already becoming a political and legal battleground. * Practical advice for users: whether it’s worth being polite to your chatbot, and what low-cost things you can do if you want to hedge against the possibility that these systems might matter morally. * Whether discourse about AI consciousness functions as hype or propaganda for AI companies, and why Rob thinks AI companies actually have an incentive to downplay AI consciousness. Links and further reading * Eleos AI Research — Rob’s nonprofit. Home to their research agenda, team page, and blog. If you want to follow the institutional effort on AI welfare, start here. They’re also, as Rob mentioned in the episode, actively fundraising and hiring. * “Taking AI Welfare Seriously” (Long, Sebo, Butlin et al., 2024) — the flagship report, co-authored with Jeff Sebo, David Chalmers, Jonathan Birch, and others. Argues that there’s a realistic near-future possibility of conscious or robustly agentic AI systems, and lays out concrete steps AI companies should be taking now. * “Consciousness in Artificial Intelligence: Insights from the Science of Consciousness” (Butlin, Long et al., 2023) — the “indicators” paper referenced several times in the episode. Surveys leading neuroscientific theories of consciousness and derives computational properties you’d look for in an AI system. S * Rob’s Substack, Experience Machines — where Rob writes more informally. The piece we discussed in the episode, “Language models are different from humans, and that’s okay,” is a good entry point, as is his “Can AI systems introspect?”. * Anthropic’s “Exploring model welfare” post — the research program under which the welfare evaluations Rob discusses were conducted. Relevant both as a primary source and as evidence that at least one major lab is treating these questions as more than an academic curiosity. * Henry’s “Consciousness, Machines, and Moral Status” — Henry’s paper arguing that debates about AI consciousness are unlikely to be settled by the science of consciousness alone, and will instead be shaped by shifts in public attitudes as social AI becomes more widespread. Closely related to the public-opinion thread toward the end of the episode. * Henry’s “All too human? Identifying and mitigating ethical risks of Social AI” — Henry’s broader survey of the ethical terrain around conversational AI systems designed for companionship, romance, and entertainment. Useful background for anyone who thinks the “AI girlfriend” phenomenon is a fringe concern. * Rob’s long conversation with Luisa Rodriguez on the 80,000 Hours podcast — a three-and-a-half-hour deep dive if you want to hear more from Rob. Transcript (Please note that this transcript was lightly AI-edited and may contain minor mistakes) Henry Shevlin: Welcome back. I’m thrilled to say that our guest today here on Conspicuous Cognition is Robert Long — or Rob, as he’s known to friends — one of the most important people thinking about AI and moral status on the planet right now. Rob is the founder of Eleos AI, a research nonprofit that, in the space of about 18 months, has dragged the question of whether AI systems might one day be moral patients from the philosophical wilderness into the boardrooms of frontier AI labs. He’s the co-author of “Taking AI Welfare Seriously,” as well as the landmark “Consciousness Indicators” paper with Patrick Butlin and other authors. Rob also conducted the first ever officially commissioned welfare evaluation of a frontier model. Before Eleos, he was at the Center for AI Safety and at the Future of Humanity Institute, and he did his PhD at NYU with Dave Chalmers. He’s also, I should say, one of my favourite interlocutors on these questions anywhere in the world, and I’ve been looking forward to this conversation for months. So Rob, welcome. Robert Long: Thanks so much, Henry. Likewise — and Dan, it’s great to meet you. I’ve been following your work. I’m really excited to talk to you about these issues. Henry: Fantastic. So for people who aren’t familiar with Eleos AI, can you tell us a little bit about what it is and how it came about? Rob: Yeah, so I guess we have been around for 18 months. When you said that number, I was like, whoa, has it really been that long? Time is just so weird when you work on AI. That was, I don’t know, a billion years in AI progress time, but also it feels like it was just last week in my personal life. Anyway — Eleos Research is a research nonprofit. We’re about four people. We work on the question of when and whether AI systems will be conscious or otherwise merit moral consideration, with a special focus on what we should do now: collectively, as a society, as AI companies, as policymakers. We think this is an extremely neglected issue. We’re building these really complicated AI systems. They kind of look like minds, but we don’t really understand their potential welfare. So we’re just trying to make progress on this and get more people to take it seriously. It got started because I was beginning to work on these issues organically — I’d worked on them as a philosopher, I’d worked on them at the Future of Humanity Institute. But Anthropic had actually approached me and some colleagues for advice on these issues. And in the first instance, I was having logistical problems hiring a team and assembling a team as an individual. Someone suggested I have my own bank account, or some way to pay people. And then Eleos kind of organically grew out of that and has now grown into a fully-fledged org in its own right. Henry: Out of interest, Rob — is there any degree to which this was motivated or informed by your personal interactions with LLMs, or was it more just the philosophy that motivated it? Was there any sort of moment where you were talking to an early Claude or ChatGPT version where you started to worry about welfare considerations? Rob: That’s a great question, and I’d be curious to hear your thoughts on this as well. I think it’s very easy to work on this and mostly be having it as arguments on a page or arguments in your head. I’m one of those people who doesn’t feel the AGI deep in my bones that often — although I do feel the AGI in an intellectual sense. But there have been a few times I’ve gotten a little spooked or jolted. One was reading the GPT-4 system card and just seeing the numbers of it, you know, passing various exams like the SAT. I remember that just really freaking me out, both from a safety perspective and a welfare perspective. The thing that made me start really viscerally feeling like we’re going to have to address this issue one way or the other was the Blake Lemoine incident. As many of your listeners might recall, Blake Lemoine was a Google engineer who blew the whistle because he came to believe he was talking to a sentient, conscious AI system. He got fired by Google for this, and then there was this huge bit of discourse — the first major bit of discourse on consciousness, sentience, moral status, and contemporary AI systems. I think it was one of the first times people started really caring what I was tweeting or what I was working on. You might have experienced a similar thing, Henry — the Blake Lemoine bump. From that moment, I have viscerally felt like: wow, this is going to get really confusing. People are certainly going to think AI systems are conscious. The future is going to be really w

    Should We Care About AI Welfare? (with Robert Long)

About

A podcast about big questions in philosophy, psychology, evolution, politics, artificial intelligence, and more. www.conspicuouscognition.com

You Might Also Like