Women in AI Research (WiAIR)

WiAIR

Women in AI Research (WiAIR) is a podcast dedicated to celebrating the remarkable contributions of female AI researchers from around the globe. Our mission is to challenge the prevailing perception that AI research is predominantly male-driven. Our goal is to empower early career researchers, especially women, to pursue their passion for AI and make an impact in this rapidly growing field. You will learn from women at different career stages, stay updated on the latest research and advancements, and hear powerful stories of overcoming obstacles and breaking stereotypes.

  1. 3d ago

    What Makes a Sentence Memorable? Inside Language, Memory and LLMs, with Dr. Greta Tuckute

    Why don't bigger LLMs look more like the human brain? In the brain's language network, today's large models explain only slightly more variance than GPT-2 XL - and the reason says a lot about what those brain regions actually do. Dr. Greta Tuckute (Research Fellow at the Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard) joins Jekaterina Novikova on Women in AI Research to unpack what brain-LLM alignment does and does not tell us. Greta works where neuroscience, cognitive science and AI meet - and when asked to keep only one of those three labels, AI is the first one she drops. The conversation covers how brain alignment develops over training, how sparse autoencoders can turn the "you are just comparing one black box to another" critique into something testable, and what memory experiments reveal about how meaning is stored - from why "pineapple" sticks in memory and "light" does not, to which sentence embeddings predict what people remember. In this episode: Brain alignment tracks formal linguistic competence, not reasoning - and it emerges after not much more than a developmentally plausible ~100M tokens, not 300BWhy better next-word prediction stops meaning "more brain-like" once a model has mastered languageSparse autoencoder features plus surprisal: for one frontal brain region, surprisal alone does almost as well as 32,000 SAE features, while some voxels are captured by just six features- Brains and LLMs share the main, high-variance features of language - not the idiosyncratic onesWhy learning from BPE tokens puts brain-LLM comparisons on "pretty shaky ground", and what should come nextA distinctive meaning makes words and sentences memorable - and SBERT predicts human sentence memory better than the other embedding models testedFrom running a photography business at 14 to a PhD at MIT, and why trying the other path first was worth it PAPERS DISCUSSED From Language to Cognition: How LLMs Outgrow the Human Language NetworkInterpreting Brain Responses to Language with Sparse Features from Language ModelsDriving and suppressing the human language network using large language models Intrinsically memorable words have unique associations with their meaningsA distinctive meaning makes a sentence memorable MENTIONED IN THIS EPISODE Anna Ivanova on formal vs functional linguistic competence (WiAIR) GRETA TUCKUTE Website: http://www.tuckute.comBluesky: https://bsky.app/profile/gretatuckute.bsky.socialX: https://x.com/GretaTuckute Women in AI Research (WiAIR) is a podcast and YouTube channel where Jekaterina Novikova talks with women doing AI research about their work, the questions driving it, and the paths that brought them there. 🎧 Subscribe to stay updated on new episodes spotlighting brilliant women shaping the future of AI. Follow WiAIR at: ⁠⁠⁠LinkedIn⁠⁠⁠⁠⁠⁠Bluesky⁠⁠⁠⁠⁠⁠X (Twitter)⁠⁠⁠⁠⁠⁠⁠WiAIR website⁠

  2. Aug 19

    Is an Image Worth a Thousand Words? Hidden Failures of Multimodal Metrics, with Dr. Elisa Kreiss

    An image is not worth a thousand words - it's worth an indefinite number of them. So why do the metrics we use to evaluate AI-generated image descriptions still assume there's one correct answer? In this episode of Women in AI Research, I talk with Elisa Kreiss (Assistant Professor of Communication at UCLA, director of the Coalas Lab) about what happens when you actually test the metrics the field relies on, and why CLIPScore, one of the most widely used measures for scoring image descriptions, stops correlating with human judgment the moment you introduce context. We also get into why longer descriptions aren't necessarily more informative, what happens when you just ask a model to "be concise," and whether AI models trip over charts and graphs the same way humans do. Elisa's research sits at the intersection of linguistics, accessibility, and multimodal AI, and this conversation covers the full arc of her work, from the theoretical question of why humans never describe images the same way twice, to the practical question of what that means for building systems that actually work for blind and low-vision users.In this episode: Why "context matters" is more radical than it sounds for image description evaluationThe hidden reason CLIPScore breaks down once context enters the pictureWhy length is a bad proxy for information density - and what to use insteadWhat happens when you prompt a model to just "be concise"Why charts and photos need completely different evaluation approachesWhether AI models make the same mistakes as humans when reading data visualizationsWhat NeurIPS's Top Reviewer Award taught her about writing a genuinely useful peer review Resources & Links: Context Matters for Image Descriptions for Accessibility: Challenges for Referenceless Evaluation MetricsWhen More Words Say Less: Decoupling Length and Specificity in Image Description EvaluationCHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models 🎧 Subscribe to stay updated on new episodes spotlighting brilliant women shaping the future of AI. Follow WiAIR at: ⁠⁠LinkedIn⁠⁠⁠⁠Bluesky⁠⁠⁠⁠X (Twitter)⁠⁠⁠⁠⁠WiAIR website⁠

  3. Jul 22

    Is Your AI Just Flattering You? Sycophancy, AI Policies, and More, with Dr. Malihe Alikhani

    Only 19% of Americans say AI has actually improved their productivity - so why the gap between the hype and reality? In this episode of Women in AI Research, Dr. Malihe Alikhani (Northeastern University, Contextual AI Lab) unpacks the hidden failures in how we build and deploy AI: why sycophancy is really a collapse of alignment, why bigger models aren't better aligned, and why "thin" alignment breaks down in the real world. Key topics The impact of moving across different AI contexts on system designThe role of language as performative and active in shaping realityInteractive inference and uncertainty in AI systemsThe importance of context in meaning and system designAI policy, transparency, and societal impactSycophantic behavior in large language modelsMeasuring AI alignment: thin vs. thickAI adoption across sectors and demographic groupsThe role of policy in AI development and safetyEthical considerations in AI research and deployment Resources & Links: Breaking the AI Mirror: Sycophancy, productivity, and the future of collaborationHype and harm: Why we must ask harder questions about AI and its alignment with human valuesHow are Americans using AI? Evidence from a nationwide survey Connect with Dr. Malihe Alikhani: https://x.com/malihealikhani 🎧 Subscribe to stay updated on new episodes spotlighting brilliant women shaping the future of AI. Follow WiAIR at: ⁠LinkedIn⁠⁠Bluesky⁠⁠X (Twitter)⁠⁠⁠WiAIR website⁠

  4. Jun 17

    Can Language Alone Create Intelligence? Insights from Neuroscience and AI, with Dr. Anna Ivanova

    Do large language models truly understand language—or are they sophisticated pattern matchers? In this conversation, Dr. Anna Ivanova (Asst. Prof. at Georgia Tech) explores one of the important questions in AI: the relationship between language, thought, and intelligence. Drawing from neuroscience, cognitive science, and AI research, Anna explains why language understanding is harder to define than most people realize, why reasoning and language are not the same thing, and what today's LLMs can and cannot tell us about human cognition. Key Topics: Do LLMs understand language or merely generate convincing text?The difference between formal and functional linguistic competenceWhat LLMs can learn from language alone—and what they cannotWhy human cognition and AI cognition may be fundamentally differentTheory of mind, reasoning, and common misconceptions about AI capabilitiesHow cognitive scientists evaluate the "thinking" abilities of LLMsWhat neuroscience can teach AI researchers about interpretabilityWhy understanding AI requires studying both behavior and internal representationsThe future of multimodal models and AI cognition Resources & Links: What does it mean to understand language?Dissociating language and thought in large language modelsHow to evaluate the cognitive abilities of LLMsHow Do LLMs Use Their Depth?True Lens Connect with Dr. Anna Ivanova: https://bsky.app/profile/neuranna.bsky.social https://x.com/neuranna 🎧 Subscribe to stay updated on new episodes spotlighting brilliant women shaping the future of AI. Follow WiAIR at: LinkedInBlueskyX (Twitter)⁠WiAIR website⁠

  5. Apr 13 ·  Bonus

    EACL 2026: LLMs Can Hear… But Can They Reason? A New Benchmark for Audio Intelligence

    What does it actually mean for a model to understand audio Paper: https://arxiv.org/abs/2601.19673 In this episode, I talk with Iwona Christop, a PhD student at Adam Mickiewicz University, about her recent EACL paper introducing ART (Audio Reasoning Tasks) — a new benchmark designed to evaluate whether multimodal LLMs can truly reason over audio, not just transcribe or classify it. Most existing benchmarks test audio skills in isolation (like ASR or classification). But real-world intelligence requires something deeper: combining signals, comparing sounds, tracking context, and making decisions. This work takes a different approach: No text-only shortcuts — tasks can’t be solved via transcription aloneReasoning-first design — models must combine multiple audio cuesNo expert knowledge required — anyone can verify correctness We also dive into the diverse task design, including: Audio arithmetic (counting and comparing sounds)Cross-recording speaker & language identificationSound-based reasoning (e.g., inferring properties from audio)Speech feature comparison (accents, variations)Multimodal reasoning across text and sound The dataset includes 9 tasks, 9,000 samples, and 30+ hours of audio — all generated in a scalable way using templates and TTS. 👉 If you care about multimodal reasoning, evaluation, or the limits of current LLM capabilities, this conversation is for you. Iwona Christop: https://www.linkedin.com/in/iwona-christop/ 👍 Like & subscribe for more deep dives into cutting-edge AI research 🔔 New episodes from EACL 2026 coming soon #WiAIR #EACL2026

About

Women in AI Research (WiAIR) is a podcast dedicated to celebrating the remarkable contributions of female AI researchers from around the globe. Our mission is to challenge the prevailing perception that AI research is predominantly male-driven. Our goal is to empower early career researchers, especially women, to pursue their passion for AI and make an impact in this rapidly growing field. You will learn from women at different career stages, stay updated on the latest research and advancements, and hear powerful stories of overcoming obstacles and breaking stereotypes.

You Might Also Like