Tried and Tested Podcast

PSI Services

Welcome to PSI’s Tried and Tested Podcast, hosted by Isabelle Gonthier. Tune in every other Thursday for expert conversations on testing, credentialing, and assessment. Each episode explores industry trends, practical insights, and the real work behind high stakes testing, featuring leaders and practitioners shaping the future of the assessment industry.

  1. Jul 23

    AI and the New Challenges Facing Educational Testing with UMass Amherst’s Stephen Sireci

    In this episode of Tried and Tested, host Isabelle Gonthier, PhD, ICE-CCP speaks with Stephen G. Sireci, Ph.D., Distinguished Professor and Executive Director of the Center for Educational Assessment at the University of Massachusetts Amherst, about AI and the new challenges facing educational testing. Drawing on his work in psychometrics, validity, fairness, educational measurement, and test development, Stephen explores how AI is reshaping what assessment professionals can do, and what they must continue to scrutinize carefully. As generative AI becomes part of test development, scoring, translation, accessibility, reporting, and candidate preparation, this conversation examines what should change and what must remain foundational. Stephen discusses how assessment can better support learning, why AI outputs still require human interrogation, and what personalization means for long-standing ideas of standardization. He also explores validity evidence in AI-supported workflows, cultural and linguistic fairness, micro-credentials, and why validity research remains essential to the future of the testing profession. What you’ll learn: How AI is expanding what assessment professionals can do across the testing lifecycleWhy assessments are increasingly being used to support learning, not only measure itWhat personalized assessment models could mean for validity, fairness, and traditional ideas of standardizationWhy AI outputs need to be interrogated, evaluated, and supported by evidenceHow AI may affect test development, scoring, alignment, translation, and accessibilityWhat candidate use of AI means for defining the construct being assessedWhere cultural and linguistic context matters when applying AI in assessmentWhy validity research remains central as testing programs adopt new technologiesWho should listen: Educational testing and assessment leaders evaluating AI-enabled changePsychometricians and measurement professionals focused on validity and fairnessTest developers exploring AI-supported item development, scoring, or alignmentCredentialing and certification leaders considering AI’s impact on workforce skillsLanguage assessment and multilingual testing professionalsResearchers and graduate students working in educational measurementOrganizations balancing innovation, defensibility, and public trust in testing About the guest:Stephen G. Sireci, Ph.D. is Distinguished Professor and Executive Director of the Center for Educational Assessment in the College of Education, University of Massachusetts Amherst. He earned his Ph.D. in psychometrics from Fordham University and his master and bachelor degrees in psychology from Loyola College Maryland. Before UMass, he was Senior Psychometrician at GED Testing Service, Psychometrician for the CPA Exam and Research Supervisor of Testing for the Newark NJ Board of Education. He is known for his research in validity and fairness of educational tests, and for innovations in test development. He currently serves/has served on several advisory boards including the National Board of Professional Teaching Standards, Duolingo English Test, and technical advisory committees for Florida, Maryland, New Hampshire, New York, Montana, Puerto Rico, and Texas. He is a Fellow of American Educational Research Association and of Division 5 of American Psychological Association, and a lifetime member of the National Academy of Education. He is a past President of International Test Commission, Northeastern Educational Research Association, and National Council on Measurement in Education. His UMass honors include School of Education’s Outstanding Teacher Award, Conti Faculty Fellowship, Public Engagement Fellowship, Outstanding Accomplishments in Research and Creative Activity Award, and the Chancellor’s Medal. He also received the Messick Memorial Lecture Award from Educational Testing Service/International Language Testing Association.

  2. Jul 9

    Global Learning, AI, and the Future of Workforce Readiness with Maria Spies of HolonIQ and QS

    In this episode of Tried and Tested, host Isabelle Gonthier, PhD, ICE-CCP, speaks with Maria Spies, Co-Founder of HolonIQ and Chief Innovation Officer at QS, about global learning, AI, and the future of workforce readiness. Drawing on more than 25 years in post-secondary education, workforce learning, and education technology, Maria explores how demographic change, technology, and shifting workforce needs are reshaping the way learners build skills and how organizations recognize them. As governments, employers, and education providers respond to rapid change, Maria explains why assessment and credentialing leaders need to think beyond traditional models of learning and proof. The conversation covers micro-credentials, alternative pathways, AI literacy, learning analytics, embedded learning in the flow of work, English language learning for global mobility, and why trusted evidence of skills will become increasingly important across the lifelong learning ecosystem. What you’ll learn: How global education and workforce trends are reshaping assessment and credentialingWhy demographic shifts are creating different learning and skills needs across regionsWhat is driving demand for micro-credentials, short courses, and alternative pathwaysWhy proof of skills is becoming more important, and more complex, for learners and employersHow technology and data can help make learning more observable without losing sight of design and trustWhat embedded learning in the flow of work could mean for future assessment modelsWhy English language learning is a useful signal for broader workforce learning trendsHow leaders can focus innovation by building from their organization’s existing strengthsWho should listen: Assessment and credentialing leaders planning for workforce changeEducation and training organizations building flexible learning pathwaysWorkforce development leaders focused on upskilling and lifelong learningEmployers and talent leaders thinking about proof of skills and workforce readinessEdtech, learning science, and product teams designing digital learning experiencesLanguage assessment and English learning professionals tracking global mobility trendsOrganizations exploring how AI, data, and credentials can support trusted skill recognition

  3. Jun 25

    From Point-in-Time Testing to Continuous Certification: Insights with NBCRNA’s Tim Muckle

    In this episode of Tried and Tested, host Isabelle Gonthier, PhD, ICE-CCP, speaks with Tim Muckle, PhD, ICE-CCP, Chief Assessment Officer at the National Board of Certification and Recertification for Nurse Anesthetists, about longitudinal assessment and the future of continued professional certification. Drawing on his background in mathematics, psychometrics, healthcare certification, and assessment leadership, Tim explains how NBCRNA is rethinking recertification in a way that supports patient safety, lifelong learning, and trust with certificants. As certification programs look for models that are more continuous, flexible, and relevant, Tim shares how NBCRNA moved from a traditional point-in-time assessment toward the MAC Check, a longitudinal assessment model built around smaller quarterly question sets, immediate feedback, spaced repetition, and ongoing evidence of competence. The conversation also explores stakeholder engagement, change management, the balance between rigor and accessibility, how AI may influence clinical practice and assessment, and why mission alignment is essential when changing maintenance of certification programs. What you’ll learn: How longitudinal assessment differs from traditional point-in-time testingWhy continuous certification can better support lifelong learning and professional growthHow smaller, periodic question sets can reduce burden while still accumulating evidence of competenceWhat immediate feedback, spaced repetition, and related retake questions add to the assessment experienceWhy psychometric defensibility alone is not enough to build stakeholder trustHow NBCRNA used research, communication, and certificant feedback to support program changeWhere AI, evolving clinical practice, and expectations for relevance may shape future certification modelsWhy assessment leaders should start with mission before redesigning recertification or maintenance programsWho should listen: Certification and recertification leaders exploring longitudinal assessmentAssessment and credentialing professionals responsible for maintenance of certification programsPsychometricians and test developers balancing rigor, feedback, and learner supportLicensure and healthcare certification organizations focused on public safety and continuing competenceProgram owners managing stakeholder trust, communication, and changeOrganizations rethinking how certification can better support lifelong professional learningAbout the guest:Dr. Muckle is the Chief Assessment Officer with the National Board of Certification and Recertification for Nurse Anesthetists (NBCRNA), based in Chicago, IL. In his career, Dr. Muckle has been responsible for overseeing the test development and psychometric quality of a variety of credentialing examinations. His work has included the development of innovative assessment formats (such as longitudinal assessment), consultation and steering of testing policy and strategy, and spearheading an assessment-focused research agenda. He brings 25+ years of experience in certification, test development and research to his present role. Tim has published numerous articles in testing-related scientific journals and publications and has been a regular presenter at meetings of the Association of Test Publishers (ATP), and the Institute for Credentialing Excellence, among others. Among his skills and qualifications are: strategic thinking and planning, organizational leadership, statistical analysis, research and publications, team building, innovative assessment formats, and volunteer service to the credentialing industry.

  4. Jun 11

    What AI Changes About Assessment Evidence: Insights with Khan Academy’s Kristen DiCerbo

    In this episode of Tried and Tested, host Isabelle Gonthier, PhD, ICE-CCP, speaks with Dr. Kristen DiCerbo, Chief Learning Officer at Khan Academy and one of Time’s top 100 people influencing the future of AI in 2024. Drawing on her background in school psychology, learning science, edtech, and assessment design, Kristen explores how AI is changing what assessment professionals can observe, interpret, and use as evidence of what learners know and can do. As assessment programs look for more authentic ways to measure real-world skills, Kristen explains why better tasks are only part of the equation. The conversation covers task models versus evidence models, the importance of closing the inferential distance, lessons from simulation and game-based assessment, Khan Academy’s “Explain Your Thinking” work, and why new measurement approaches should often begin in lower-stakes formative environments before moving into higher-stakes use. What you’ll learn: How AI can surface evidence of learner thinking beyond a single answerWhy valid assessment starts with separating tasks from evidenceWhat “closing the inferential distance” means for assessment designWhere simulations, games, and open-ended tasks can create noise instead of signalWhy new AI measurement approaches should start in low-stakes settingsHow to balance innovation with structure, validity, and defensibilityWhat multimodal AI and on-device models could mean for future assessmentWho should listen: Assessment and credentialing leaders exploring AIPsychometricians and test developersCertification and licensure program ownersEdtech and learning science teamsProduct leaders building digital assessment experiencesEducators working with formative or performance-based assessmentAbout the guest:Dr. Kristen DiCerbo is the Chief Learning Officer at Khan Academy, where she leads the content, assessment, design, product management, and community support teams. Time magazine named her one of the top 100 people influencing the future of AI in 2024. Dr. DiCerbo’s career has focused on embedding insights from education research into digital learning experiences. Prior to her role at Khan Academy, she was Vice-President of Learning Research and Design at Pearson, served as a research scientist supporting the Cisco Networking Academies, and worked as a school psychologist. Kristen has a Ph.D. in Educational Psychology from Arizona State University.

  5. May 28

    Test Security, Agentic AI, and the Future of Assessment with Paul Muir from risr/

    As assessment programs confront increasingly sophisticated fraud and rapid advances in Artificial Intelligence, protecting trust and integrity has never been more complex. In this episode of Tried and Tested, host Isabelle Gonthier, PhD, ICE-CCP sits down with Paul Muir, Chief Customer Officer at risr/ and Board Chair of the Association of Test Publishers, for a candid conversation recorded live at the ATP Innovations in Testing Conference. With more than 25 years in the assessment industry, Paul shares a front‑line perspective on test security, the rise of agentic AI, and the evolving role of industry collaboration. Together, they explore how assessment leaders can move from reactive security measures to more strategic, forward‑looking approaches, and how ATP is helping guide the community through a period of rapid technological change. What you’ll learn: How the test security threat landscape is evolving, including organized cheating, content harvesting, and increasingly sophisticated fraud techniquesWhere assessment and credentialing programs remain most vulnerable today, and why keeping pace with fraud is such a challengeWhat agentic AI means in the context of assessment, and how more autonomous systems introduce both new risks and new opportunitiesHow AI can be used proactively to strengthen test security rather than simply reacting to emerging threatsHow Paul’s dual perspective as a technology leader and ATP Board Chair informs his view of where assessment is headedThe role ATP plays in helping the industry navigate innovation, security, and trust during periods of rapid changeWhy collaboration and community engagement are critical to protecting the integrity of assessment programsWho should listen: Assessment and credentialing leaders responsible for program integrityTest security, compliance, and risk professionalsEducation and assessment technology providersOrganizations exploring or deploying AI in assessment programsATP members and professionals engaged in the broader assessment communityAbout the guest:Paul Muir serves as the Chief Customer Officer at risr/ and currently holds the position of Board Chair for the Association of Test Publishers (ATP). With over 25 years of experience in the assessment sector, Paul is an experienced leader and active volunteer, specialising in areas such as test security, assessment reform, technology-enabled assessment, and Artificial Intelligence. Since joining risr/ in 2024, Paul, who is based in the UK, has been responsible for driving thought leadership, developing strategic partnerships, and fostering engagement across the industry, community, and customer base.

  6. May 14

    Serious Games and Authentic Assessment with Jenn McNamara from BreakAway Games

    In this episode of Tried and Tested, host Isabelle Gonthier, PhD,- ICE CCP, Chief Assessment Officer at PSI and ETS, sits down with Jenn McNamara, Vice President of Strategic Products and Partnerships at BreakAway Games. Recorded live at the ATP Conference, the conversation explores how serious games are being used to support authentic, performance based assessment across defense, healthcare, and professional credentialing programs. Jenn shares how game based environments can capture real world decision making, reduce traditional test effects, and provide richer evidence of readiness than item based assessments alone. The discussion also covers the role of AI in game design and assessment delivery, the importance of accessibility by design, and what emerging innovation trends signal about the future of assessment. What you’ll learn: How “serious games” are defined and how they differ from surface level gamificationWhy authenticity, particularly cognitive fidelity, matters more than visual realism in assessment designWhat immersive, strategy based environments can reveal about real world decision making, stress, and readinessHow independent research and validation help make game based assessment credible and defensible in high stakes contextsWhere serious games are most effective in assessment programs, and where assessment leaders should be cautiousHow AI is influencing the future of immersive assessment, including adaptive scenarios and content variation, while maintaining rigor and fairnessWho should listen: Credentialing and certification leaders exploring authentic or performance based assessment approachesAssessment professionals seeking new ways to evaluate decision making, judgment, and applied skillsPsychometricians and assessment designers interested in emerging item types and validation modelsProgram owners responsible for accessibility, candidate experience, and exam integrityTesting and education technology professionals tracking the role of serious games and AI in assessment About the Guest: Jenn McNamara is Vice President of Strategic Products and Partnerships at BreakAway Games, focused on delivering immersive solutions for education, assessment, and performance support. With over 25 years of experience in cognitive psychology, serious games, and AI-driven systems research and development, Jenn is a recognized leader and trusted partner raising the standard of advanced, technology-driven learning and assessment for defense, healthcare, and corporate sectors. Jenn’s pioneering design and development approaches set the accessibility standard for serious games. She speaks frequently at conferences including I/ITSEC, ATP Innovations in Testing, Serious Play, and Games for Change. In volunteer support of the industry, Jenn also directs the nonprofit Serious Games Showcase & Challenge and serves in leadership roles with the NDIA Human Systems Conference, Serious Play Conference, and ATP Innovations in Testing Conference.

  7. Apr 30

    Libby Rodney from The Harris Poll on the 2026 ETS Human Progress Report: Adaptability, Credentials & Opportunity

    This is our 50th episode of Tried & Tested and we’re marking the milestone with a conversation about what may be one of the biggest questions facing the workforce right now: how do people prove what they can do as change accelerates? Host Isabelle Gonthier welcomes Libby Rodney, Chief Strategy Officer at The Harris Poll, to unpack what “proof” looks like in today’s disrupted environment and why credentials are increasingly tied to confidence, mobility, and opportunity. Drawing on The Harris Poll research behind the 2026 ETS Human Progress Report, Libby Rodney explains how leaders can move beyond noisy narratives and toward signals that actually matter, especially as workers face rapid disruption, shifting skill demands, and rising pressure to demonstrate adaptability. Together, they explore the widening gap between interest and access, the need for clearer employer signals about what’s valued, and why the future of credentialing depends on trust, alignment, and measurable evidence, not just storytelling. What you’ll learn: What “proof of skills” really means in a labor market shaped by constant disruption and why workers are seeking evidence they can carry across roles and industries.How the 2026 ETS Human Progress Report was designed, including the role of benchmarking and how the focus evolved toward adaptability as a cornerstone theme.Why credentials are rising in importance right now, including the connection between workplace anxiety and the need for verifiable evidence of capability.What’s driving the demand vs. access gap (high interest, lower access), and the practical barriers that keep people from pursuing credentials.Why employer clarity is the unlock: what workers need employers to specify about which credentials matter, and why ambiguity discourages investment before cost even enters the picture.How AI is reshaping credential expectations, including why many workers want formal certification to verify AI skills, and why “using AI” isn’t the same as using it well.Who should listen: Credentialing, certification, and assessment leaders designing programs that must earn trust and prove value.Employers, HR, and workforce strategy teams deciding which skills and credentials matter most.Education and training leaders working to close the gap between learning pathways and real-world opportunity.About the guest: A scenario planner, cultural strategist, and navigation expert, Libby Rodney serves as Chief Strategy Officer at The Harris Poll where she helps Fortune 100 executives decode uncertainty and navigate transformational change.  Creator of “The Next Big Think!” substack and co-host of “So Get This” podcast (new on Bubbler/iHeart Radio), her cultural intelligence framework has predicted major shifts including “quiet vacationing,” FOBO (fear of being obsolete), and the “lottery over logic economy.”  Libby has commanded global stages at Davos, Cannes Lions, SXSW, Forbes CMO summit, Ad Week, and CES, establishing her as the go-to cultural decoder for organizations seeking to see around corners.  Libby’s insights have been featured across major media outlets where she has become known for revealing not just what’s trending, but what companies must pay attention to in the next 18 months.

  8. Apr 16

    Closing the Modality Gap: How TOEFL Built Trust in Remote Testing

    As questions around remote testing and test security continue to surface across the assessment landscape, how can programs move beyond perception and focus on evidence? In this episode of Tried and Tested, host Isabelle Gonthier is joined by Paul Gollash, SVP of TOEFL and GRE at ETS, and Wally Dalrymple, Chief Security Officer at PSI and ETS, for a timely conversation on trust, data, and delivery modalities. Together, they explore how the TOEFL program has evolved a layered security model across both test center and remote delivery. The result is measurable outcomes showing that differences between remote and test center delivery have narrowed to a minimal level. From identity verification and fraud indicators to data driven decision making and continuous improvement, this episode offers a practical and credible look at how security, scale, and trust come together in modern assessment programs. What you’ll learn: Why remote testing and test center delivery have distinct risk profiles, and how purpose built security approaches for each modality strengthen overall test integrityHow the TOEFL program designed and evolved a layered security model that spans the full assessment lifecycle, from registration and identity verification through delivery, review, and continuous improvementWhat layered security looks like in practice, including how multiple controls work together to deter fraud, detect risk, and protect score integrity when individual signals alone are insufficientHow ETS and PSI measure the real-world impact of specific security controls, including what the data reveals when individual layers are modified or removedWhy outcome based measures, such as score patterns and pass rate alignment, provide a more meaningful indicator of trust than incident counts or flagged events aloneHow operating at global scale enables stronger pattern detection, faster response to emerging threats, and continuous refinement of security strategies across programs and modalitiesHow AI is influencing both sides of the equation, accelerating new forms of fraud while also strengthening detection, analysis, and decision-making within modern test security programsWhat assessment leaders should consider as trust, security, and delivery models continue to evolve in an environment where risk, technology, and stakeholder expectations are constantly changingWho should listen: Assessment, credentialing, and certification leaders responsible for program integrity and long term trustTest security, fraud prevention, and risk professionals designing or evaluating security models across delivery modalitiesHigher education institutions, regulators, and score accepting organizations seeking evidence based perspectives on remote testing outcomesProgram owners and product leaders balancing access, candidate experience, and rigorous security requirementsOrganizations navigating concern around the credibility of remote English language testing and other high stakes assessments

Ratings & Reviews

5
out of 5
4 Ratings

About

Welcome to PSI’s Tried and Tested Podcast, hosted by Isabelle Gonthier. Tune in every other Thursday for expert conversations on testing, credentialing, and assessment. Each episode explores industry trends, practical insights, and the real work behind high stakes testing, featuring leaders and practitioners shaping the future of the assessment industry.