The Augmented Educator Podcast

Michael G Wagner

Stories From Education's AI Frontier. Exploring the promises, pitfalls, and possibilities of algorithmic teaching and learning. An AI-voiced companion to the Substack of the same name. www.theaugmentededucator.com

  1. 6h ago

    Nobody Wept for the Weavers

    I’ve held off on writing this post for some time. The subject takes us outside education, and my position on it is unconventional enough that I kept finding reasons to postpone. Neither is much of an excuse though. And the detour may turn out to be more relevant for education than I had originally thought. Creative exceptionalism is the belief that of everything human beings do, creative and knowledge work are more valuable. They are presumed to be closer to the very essence of what makes us human. Under this view, it is unfortunate but acceptable when automation pushes a tradesperson onto an automated factory floor. But when the same automation arrives for a copywriter or an illustrator, it becomes an assault on the human spirit. A fair number of people I read and respect hold some version of this belief, and at least some of them are likely going to be annoyed with this post. Maybe even offended by it. But I could never really make sense of this worldview. I have always felt it is somewhat elitist. Most of us buy mass-manufactured clothing without a flicker of guilt, but few of us have ever wondered what happened to the tailors that used to handcraft garments. At the same time, we now turn around and insist that pictures and sentences have to come from human hands, or we are losing something central to the human experience. What really separates the illustrator from the tailor? What gives one of them a claim on humanness that the other is apparently not allowed to make? So in this post, I want to analyze that question with the help of two past examples of automation-driven labor displacement. Both examples used to be discussed in a way that strongly resembles today’s argument about creative uses of AI, and both have since largely faded from our collective memory. A maiden with a pen, a harlot in print Before movable type reached Europe, copying a text was skilled and sacred work. In monastic scriptoria and secular scribal guilds, manuscript copying was understood as devotio, labor in which the hand of the scribe served as a physical conduit for divine and classical wisdom. Hand-copying and the careful curation of manuscripts were both treated as safeguards against textual corruption and against the moral decay assumed to follow from it. Then the press arrived, and the scribal elite fought it. Filippo de Strata, a Dominican friar from Venice, formally protested to Doge Nicolò Marcello, who governed Venice from August 1473 to December 1474, calling for the removal of all printing presses from the Republic. In Martin Lowry’s translation, de Strata’s most quoted line sets the two crafts against each other: “She is a maiden with a pen, a harlot in print.” The writing profession, he wrote, had been degraded in the brothel of the printing presses. Mechanical reproduction stripped writing of its dignity and let uneducated commoners reach texts with no learned intermediary standing between them and the page. Typographical quality was secondary. What de Strata objected to was who could now read what, and without whose supervision. In 1492, the Benedictine abbot Johannes Trithemius wrote De laude scriptorum, “In Praise of Scribes,” arguing that hand-copying carried a spiritual utility no press could match. The physical act of writing, he claimed, pressed sacred words deeper into the mind of the copyist, while printed books were cheap and spiritually empty. He urged his monks to keep copying. Then he had the book printed, in 1494. I find that hypocrisy almost too perfect. His case for the irreplaceable human hand was set in type on the machine it warned against, by a man who used that machine freely for his own work. If you have read much of the current commentary on AI and writing, that pattern will likely sound familiar. The loom was somebody’s identity too Three centuries later, the same argument reappeared. This time it came from people with no ecclesiastical standing at all. Handloom weaving was one of the largest working-class trades in Europe. It had a unique culture. Weavers were skilled artisans with actual control over their own hours and their own techniques. They worked inside what the social historian E. P. Thompson called a moral economy, where transactions were governed by customary obligation and community norms rather than by market logic. Weaving was tied to personal independence. It also carried the pride of taking a piece of work through from raw yarn to finished cloth. The power loom and factory system dismantled that. Mechanization severed the weaver from the finished object and reduced an autonomous artisan to an appendage of a machine. Physical technique and aesthetic judgment were pulled out of human hands and embedded into mechanical components. Luddism, between 1811 and 1816, is usually remembered as blind rage against machinery. But the Luddites were considerably more selective than that. The croppers and weavers specifically went after the machines being used to cut wages and put the weaver on the factory clock, leaving other machinery alone. Thompson’s The Making of the English Working Class, published in 1963, documents what came next: violent dislocation and the destruction of a self-governing cultural life that had taken a century to build. Compare the weaver’s argument with that of today’s illustrators. It is correct that illustrators have one grievance the weavers never had, since nobody trained a power loom on their own cloth. I have addressed that question separately in Is AI Training Stealing? But the complaint itself has the same shape. The machine cannot do what my hands do. The cloth was never the only thing being made. My craft is part of who I am. The weavers were right about most of it, but it did not matter anyway. The political economists and industrialists of the period dismissed the argument as irrational resistance to economic necessity. Human craft needed to yield to industrial efficiency. No one with the power to decide anything agreed that their kind of humanness counted. The wrong things got automated first The twentieth century then built an entire social story about which work was considered safe. Richard Florida’s The Rise of the Creative Class, published in 2002, became the best-known expression of it. Creative and symbolic work sat at the apex of human economic activity. By contrast, manual trades and routine service work were treated as inferior and inherently automatable. Education policy told people to climb out of the manual trades and into the cognitive professions. “Brain over muscle” hardened into an explicit instruction. The unspoken agreement was that physical tasks could be automated, while abstract creative work would remain human. The reality, however, is that this creative sanctuary was always an illusion, a fact that roboticists had understood since the 1980s. It became known as Moravec’s paradox. Tasks we experience as intellectual turn out to be easier to automate than the sensorimotor work we perform without conscious thought. Abstract logic and rule-governed writing are relatively manageable. Physical work in an unpredictable environment, with all the perception and touch it demands, is far harder to automate. Hans Moravec put it plainly in Mind Children in 1988, writing that it is “comparatively easy to make computers exhibit adult level performance on intelligence tests or playing checkers, and difficult or impossible to give them the skills of a one-year-old when it comes to perception and mobility.” Generative AI exposed this illusion. It reached symbolic and visual work well before it reached the plumber or the carpenter. And it caused panic the moment the automation reached the occupations everybody had been told were safe. Learn to code There is an inherent class dynamic here, and it becomes clear as soon as we compare how these labor displacements are being discussed in public. The Federal Reserve Bank of New York puts the number of United States manufacturing jobs lost between 2000 and 2010 at 5.7 million, concentrated in communities with few replacement jobs. The public narrative pointed to market efficiency and global competition, with a strong side order of personal responsibility. The public answer was retraining. Workers were expected to move into cognitive jobs, and that expectation morphed into a slogan that is now a cultural artifact in its own right: “learn to code.” And it was not only a slogan. Mined Minds was a nonprofit founded by Amanda Laucher that ran free coding bootcamps for laid-off miners in West Virginia and Pennsylvania. From around 2016 it drew national coverage as a model of what retraining could do. In May 2019, roughly sixty former West Virginia students filed a class action against it. They alleged that the promised stipends, apprenticeships, and job placements never arrived, and that they had put their lives on hold on the strength of those promises. Laucher said the filing had no merit and pointed to a student contract stating that no job or further training was guaranteed upon completion. Mined Minds became a national story about retraining as redemption. But the allegations made that story a good deal harder to defend. The prevailing language throughout was economic adjustment. The machinist’s embodied knowledge never sat at the center of that discussion. The advice was to climb the ladder into the safe work, but the safety of that work was only assumed and never guaranteed. Then the same expectation was dropped onto graphic designers, copywriters, and junior developers, in the form of “learn to prompt.” This time it produced a backlash. I do believe some anger is justified. Telling someone whose job just vanished to “learn to prompt” is dismissive. But this way of dealing with job displacement did not become dismissive when generative AI appeared. It was always dismissive. It is just that most creatives and knowledge workers had no particular reason t

  2. Sep 1

    The Noise Floor

    We need to talk about AI detection again. I know, I have been writing about it a lot lately. But the problem keeps widening, and this time the failure looks different from anything I have covered on The Augmented Educator before. In my last post, I argued against AI text detection on mathematical grounds. Human writing and machine writing produce overlapping score distributions. Some AI-generated passages will always score as more human than some human-written ones, and vice versa. Any educator who uses a detector therefore has to set a threshold somewhere inside that overlap. And wherever that dial lands, somebody always pays for it. The two distributions do not permit a clean separation. No threshold can create one. I called the piece There Is Only One Dial and hoped it would close the case. But, as I came to realize, it closed only a quarter of it. The other three quarters concern AI-generated images, video, and audio. Education has barely discussed these systems, mostly because they usually operate outside the educator’s view: in stock library upload queues, music distribution pipelines, journal submission portals, and platform moderation systems. Photographers and musicians have been arguing about these tools for years. Most educators have not heard of them. They should. Because I think these detectors may pose an even greater risk to our students than the text classifiers we have spent so much time debating. The D’Addario commercial This piece began, as many of mine do lately, with a YouTube controversy. This time, the trigger was a commercial containing a piece of music that was accused of being AI-generated. D’Addario is a family-run American manufacturer of music accessories and one of the best-known names in guitar strings. If you play, you have almost certainly had a set of theirs on an instrument. In July 2026, the company launched two extended-range string lines and posted a promotional video built around a high-gain progressive metal track. Within hours, the comments were filled with accusations. The performance had an odd digital sheen, and the timing was quantized so tightly that it no longer felt played. It lacked the microdynamics that make a guitar sound like a guitar. The accusation was clear: D’Addario had skipped the musicians and typed a prompt into the AI music generator Suno. What followed made it worse. D’Addario deleted comments, blocked accounts, and eventually switched comments off completely. The company later explained that the employee who made the track had been doxxed and was being harassed. The moderation was meant to protect him. Whatever the intent, an audience that already suspected a cover-up read the silence as confirmation. Then a behind-the-scenes video of the project session, posted by D’Addario as proof, drifted out of sync with the commercial, which viewers took as further evidence of fakery. But that assessment turned out to be wrong. Rhett Shull, a guitarist and producer whose YouTube channel covers gear for a large audience of players, obtained the original Logic Pro session from the company and audited it. The project contained real recorded guitar performances and hand-programmed MIDI. No text prompt ever wrote that song. What the session did contain, however, was a production chain pushed until it broke: pitch-corrected direct-input guitars, drum compressors stacked in series, every MIDI velocity pinned at maximum, and a dozen synthesizers packed into the midrange. And then, after export, the D’Addario employee added several rounds of automated mastering. Rhett Shull and Steve-san Onotera, another guitarist and YouTuber who posts as samuraiguitarist, both suspected that one of those mastering passes went through Suno. Mastering is the final stage of music production, when an engineer makes the adjustments that prepare a song for release. A skilled mastering engineer can make a mix sound louder, clearer, and more coherent without ever drawing attention to the work. And some of that work is now handled by automated services such as LANDR and by mastering assistants built into digital audio workstations. What is usually not known is that Suno’s approach to mastering works differently from these automated systems. Conventional automated mastering analyzes a finished stereo file, then applies equalization, compression, and limiting to it. Those are the same operations a human engineer would reach for, selected and adjusted by a model. Suno’s version, on the other hand, runs the audio back through its generative engine and completely rebuilds it. The result is a new waveform rather than a processed copy. And the consequence of that is that the finished waveform is machine-generated, even though the performance underneath it is not. I find that explanation convincing as an account of why D’Addario’s track sounded artificial, but it remains an educated guess. At the time I am writing this piece, D’Addario has not confirmed which tools were used at that stage, and nothing in the published forensic analysis can prove it either way. Part of the reason that question is still open is the underlying systemic failure Onotera pointed to. D’Addario’s communications staff were defending a technical claim they did not understand, with evidence they could not evaluate. And the music professionals watching them did not necessarily understand generative AI any better. The company needed over a week of internal investigation before it could describe its own production chain with reasonable accuracy. Hold on to that diagnosis. Of all the elements in this story, it is the one with lasting significance. Thirty-six percent of nothing In the same video, Onotera described an experiment that should worry any creator deeply. He took a track from his own catalog: human-composed, human-performed, conventionally recorded, with no generative anything anywhere in the chain. He then uploaded it to the AHA Music AI detector, an online tool used across music distribution and content monitoring workflows. The result was 36 percent AI-generated, with 95 percent confidence. For anyone familiar with the limitations of AI detection, this number should raise a big red flag. Then he ran the disputed D’Addario track through the same tool. The result was 74.7 percent AI-generated, with 95 percent confidence. Other detectors gave the same file a “Suno match” score of between 73.88 and 91.40 percent. If the mastering suspicion is right, those tools were picking up a genuine Suno signature sitting on top of a human performance. They may have been correct about the artifact, but they were wrong about everything anybody should ever really care about. Let’s look at those numbers more closely, starting with that confidence figure, because it is the part people misread the most. It expresses how certain the model is within its own system. It does not independently validate anything. Onotera’s control experiment shows why that distinction is worth making: the tool was highly confident about a result we know was wrong. A confidently wrong number is more dangerous than an openly uncertain one, because it gives a guess the authority of a measurement. There is also something hidden in the percentage itself. Audio engineers have a word for it. The noise floor is the bed of hiss underneath a recording, and once a signal falls below it, pulling the two apart gets difficult. In Onotera’s test, 36 percent on a track with no generative involvement at all behaved like the detector’s noise floor. More than a third of the scale had already been used up before any meaningful measurement was taken. A single control track is not a proper statistical sample. But it does show what that number is not. It is not a literal estimate of how much of a recording was generated with AI. The D’Addario track scored higher, and the tool could not tell anyone what it had found. Generated composition? Generated performance? An automated mastering pass? Or simply the artifacts of very heavy production? The scale collapses all of those into one figure and hands you a percentage. That is a different problem from the one in my last post, but it is arriving at the same place. There, two overlapping distributions meant no threshold could cleanly separate them. Here, the scale is not measuring what people think it does. Two ways to be wrong Research shows that text detectors and media detectors both fail. But they fail in opposite directions, and they take down opposite people on the way. A policy written for one of them will therefore be wrong about the other. Text is made of discrete tokens. A language model produces words by drawing the next one from a probability distribution, and a detector measures how surprising the resulting sequence is to a reference model. That is perplexity. Alongside it sits burstiness, the variation in sentence length and structure across a passage. Machine text tends to run smooth and evenly paced. Human text tends to be less steady. Those are two of the signals most text detectors lean on. The failure mode falls straight out of this mechanism. Careful, plain, grammatically conservative writing scores as machine writing. The detectors punish clarity, and they punish it hardest in the people who worked hardest to achieve it. I have written about this problem at length in previous posts on this Substack. AI images, video, and audio are different mathematical animals. They are continuous, high-dimensional signals, and most current generators build them through diffusion. Generation starts with static noise and strips it away step by step until a picture or a waveform emerges. Detectors for this kind of content hunt for the traces that denoising leaves behind. Those include reconstruction errors, frequency-domain artifacts, and spectrogram phase relationships that a physical microphone is unlikely to produce. But those traces are fragile. On the GenImage benchmark, which tests det

  3. Aug 25

    There Is Only One Dial

    Every company selling AI detection to educational institutions leads with a number. Turnitin, GPTZero, Copyleaks, and Pangram all publish one, usually an accuracy rate or a false-positive rate, often reported to two decimal places. Pangram, for example, self-reports a false-positive rate of 0.19 percent and a false-negative rate of 1.4 percent on standard datasets. A false-positive rate of 0.19 percent means that out of 1,000 human-written essays, the detector falsely flags only about two texts as AI-generated. Sounds great, doesn’t it? But unfortunately, none of those numbers can really tell you what will happen in your classroom. Because outside of a controlled lab environment, they have very little meaning. I have made this point on The Augmented Educator Substack for over two years now, particularly in The Paradox of AI Detection. But so far I have mostly asserted rather than explained what might be the most complex part of this argument: that the trade-off at the center of these tools is subject to a hard mathematical limit. Better engineering can shift that trade-off. But nothing can eliminate it. Absolutely nothing. The math is baked in. So in this post, I want to complete my core argument against AI detection. And I have decided to do it in plain language and with no complex math notation. I wanted to keep this approachable to a lay audience. If you are interested in the finer details of this argument, you might want to check out the audio deep dive podcast episode linked at the end of this post. How a score becomes a verdict Let’s start with what a detector actually returns in practice. Underneath the percentages and the categories, it is fundamentally a score. The detector assigns a text passage a number between zero and one, say 0.83 or 0.11, showing how strongly the system associates it with machine-written text. Early detectors often used that score to return a simple yes/no verdict, which ended up being highly problematic because it outsourced the integrity decision from the instructor to the detector. Most of the products available today therefore stop short of that. They present a percentage or a category and leave the final judgment of what that means to whoever reads the report. This separates the detector’s result from any misconduct accusation. It is now the teacher who makes that call and not the AI detector. Turnitin, for example, reports what share of a document it estimates was AI-written. It simultaneously cautions that the figure cannot be used as proof of misconduct. GPTZero splits a submission into percentages of AI, mixed, and human. And Pangram sorts text into human, AI-edited, and fully AI-generated, and its EditLens model estimates how much AI editing a passage received. None of that is an accusation. Every one of those reports leaves somebody else to decide what happens next, and that decision is a cutoff. Somewhere between an innocent detection report reading some percentage and an email accusing a student of misconduct, a line gets drawn. Statisticians call that line a threshold. Think of it as a dial which the teacher sets, and think of the number it points at as a bar the score has to clear before anyone acts on it. Turn the dial up, and fewer texts clear the bar. The same detector, on the same essay, leads to entirely different outcomes depending on what number that dial points at. When a company says its tool is 99.85 percent accurate, it is describing the result of applying a chosen threshold to a chosen test set. Move the dial or change the test set, and the advertised accuracy and error rates move with it. Two piles that overlap Let’s assume we want to use a detector to score an extensive set of known examples. Take thousands of texts you know a human wrote, for example, and score them all. Then take thousands you know came from a language model and score those too. You get two piles of scores. If those piles sat in separate ranges, say, with every human text below 0.4 and every machine text above 0.6, detection would be a solved problem. You would just have to put the dial in the empty gap between 0.4 and 0.6 and the system would never make a mistake in either direction. But the problem is that they do not sit in separate ranges. They never do. In 2023, Dalalah and colleagues found a substantial overlap between the score distributions of genuine and AI-generated academic writing. Later studies have found the same underlying overlap. Now consider who ends up in that overlap. Human writing lands in the machine range when it is clean and grammatically conservative. That includes a student writing carefully in a second language. It also includes technical and scientific prose, where uniformity is often a virtue, and short answers, which simply do not give the detector enough to work with. Jung and colleagues took that last problem seriously enough that in 2025 they built group-adaptive thresholds to stop short texts from being flagged at inflated rates. Machine writing, on the other hand, lands in the human range when it is uneven, or when somebody with an understanding of human writing has edited it. Once the two piles overlap, the trade-off is unavoidable. Wherever the dial sits, some human texts fall above it and get called machine-generated. And some machine texts fall below it and get identified as human. You cannot eliminate both errors simultaneously because both arise from the same overlapping region. Raise the dial to rescue the honest students, and you miss the AI texts sitting beside them. Lower it to catch those AI texts, and you will take the honest students with it. And this is not a flaw in anybody’s product. No clever design and no ingenious engineering will ever escape it, because the trade-off belongs to the statistical overlap and not to the software. Which mistake a school will tolerate A school, or any educational institution for that matter, has to choose which of the two errors it is more willing to accept. But the problem is that these errors do not carry equal consequences. Missing an AI-written essay compromises an assessment. This is most likely not a big deal. Falsely accusing a student, however, can derail a degree and follow that student for years. The consequence is that tools that flag innocent students get switched off because their vendors and the institutions that deploy them would otherwise risk being sued. UCLA and the University of Pittsburgh disabled Turnitin’s AI detection early on. Vanderbilt did the same in August 2023. And OpenAI famously shut down its own detector a month earlier, having shipped a tool that caught only about 26 percent of AI text while falsely flagging about 9 percent of human text. Flagging innocent students is not an option. And so the dial goes up. It has to, as long as language models continue to improve. And that changes which numbers are really relevant. A vendor’s advertised accuracy comes from whatever threshold the vendor picked for its own benchmark. A school cannot use that same setting. It has to turn the dial up until the tool wrongly flags only, say, one honest essay in every hundred. And it then has to ask how much AI writing it still catches once the dial is that high. In 2024, Tufts and colleagues argued that this was the only deployment metric worth reporting. They then tested popular detectors on text from models and subject areas the tools had not seen, using the kinds of prompts a curious student might try. Held to that one-in-a-hundred limit, some of those detectors caught nothing at all. Zero percent. Now, you might object that these tools were simply not built to catch clever students. So let’s look at one that was. In 2025, Lekkala’s group tested a detector that had been trained specifically on AI text which someone had deliberately reworded to hide its origin. This system already knew what a disguised passage looks like. And yet, held to that same one-in-a-hundred limit, it caught only 48.8 percent of them. So the safer the setting is for honest students, the less AI writing the detector is able to catch. And the less a dishonest student has to do to slip underneath it. The ones who still get caught at that setting are just the ones who did the least to hide. Nobody writes down where the dial sits Now, handing that decision back to the teacher sounds like the responsible thing to do, and in a sense, it is. Read the fine print on any of these products and you will find a version of the same warning. The score is not proof. It should not be the sole basis for an integrity finding. Every word of that is correct. But look at what it actually does to the numbers. The vendor calculates its accuracy figure at a threshold the vendor selected, publishes it, and then tells you not to rely on it. Meanwhile, the threshold that decides whether a student gets an accusatory email is the one being set in your classroom, by you, and nobody has ever measured what that threshold does in practice. And it never gets written down. Two instructors in the same department can look at the same 42 percent and draw the line in completely different places. And the same instructor can move it from one semester to the next without even noticing the change. This is the part I find hardest to defend. A documented number becomes an undocumented disciplinary judgment, made at the end of a long grading day, about a student whose other writing the reader may barely remember. The trade-off between the two errors did not go away when the product stopped short of a verdict. It simply moved somewhere nobody keeps records. The pile students can move So far, I have treated the two piles as fixed. But they are not. Evasion creates a second problem, and it makes the trade-off even worse. Many published benchmark results compare human writing against untouched model output. Someone pasted a prompt into ChatGPT, took the result, and scored it without changing a word. A student trying to avoid detection is highl

  4. Aug 22

    Everything on One Machine: Unsloth your AI

    I have written on this Substack about open-source models before, and today, I am going to write about it again. This is because I think it is critically important for educators to understand that AI is a general technological shift in how we interact with computers, and that the big foundational language models are only one small part of it. The other part, the one that gets considerably less attention, is the open-source ecosystem of models that anyone can download and run on their own hardware. The claim I keep making is this. If every foundational model disappeared tomorrow, the technology itself would not go anywhere. And last weeks, two releases made that claim harder to argue against. The first is Qwen3.8, an open-source model whose 27B version (27B stands for 27 billion parameters) can run on an advanced home computer setup and handle the kind of work you would normally send to a cloud service. The second is Unsloth Desktop, a free, open-source desktop application that makes running and training that model, and others, as simple as opening a document. Neither is enough on its own. A capable model buried in a Python script is not a replacement for the chat window you are used to. And a friendly interface pointed at a weak model is just a chat box with extra steps. But put them together, and you have something that actually works. I am running that combination right now on a MacBook Pro with 128 gigabytes of shared memory and it works remarkably well. So in today’s free bonus post, I want to walk through what a fully local, fully open-source AI environment looks like in practice, and explain why I think it is a preview of how we will interact with AI in the future. The entire process behind this post, from the initial research through the drafting to the final editing, ran through Unsloth Desktop with Qwen3.8 doing the model work. The only tool I kept from my old workflow is ProWritingAid for copyediting, and this runs locally on my machine as well. I did not use Claude, ChatGPT, Gemini, or any other foundational model anywhere in this process. Everything happened on my machine, and no data ever left the computer. How a fine-tuning library became a desktop app Unsloth started in 2023 as a passion project by two brothers, Daniel Han and Michael Han, in San Francisco. Daniel has a background at NVIDIA and leads the technical side. Michael handles design and product engineering. Together they built an open-source Python library for fine-tuning large language models, which they released on GitHub in December of 2023. The library solved a real problem. Fine-tuning a model, the process of training it further on your own data so it does a specific job better, used to require serious hardware and a fair amount of coding. Unsloth’s core contribution is a set of custom kernels, written in a language called Triton, that make the training run faster and use less memory. In their standard benchmarks, fine-tuning is about twice as fast and uses roughly 70 percent less video memory than the standard pipeline. That is the difference between needing a data center and being able to do the work on a single graphics card, or in some cases, on a laptop. The project gained traction quickly. It went through the GitHub Accelerator program in 2024, then the Y Combinator summer batch that same year. The team is still small, currently eight people, and they have so far raised half a million dollars across one round. The library’s repository has so far collected over 73,000 stars on GitHub, which in open-source terms is a clear sign that many people actually use it. Unsloth has since published compressed versions of many popular models on Hugging Face, including the Qwen3.8 model I am running. And they even pushed some of their fixes back into OpenAI’s own open-source model repository. Two front ends on the same engine By early 2026, the group had built two front ends on top of their library. The first, Unsloth Studio, launched in beta around March of this year. It is a no-code web interface for training, running, and exporting open models. You do not need to write a single line of code. You just load a model, point it at your data (it can build the training dataset automatically from a PDF, a CSV, a Word document, or a text file), and start a training run. Everything fits on one screen, which is a deliberate break from the notebook-and-script workflow that the original library required. Unsloth Desktop is the newer and more ambitious project of the two. It is a native application for macOS, Windows, and Linux, built on a framework called Tauri. It is also free and open source, and it runs entirely on your machine with no telemetry and no requirement to be online. The headline on the company’s website is blunt: the first desktop app to run and train AI models, open source, free, 100 percent local. Unsloth Desktop primarily runs large language models for text, but it can also run diffusion models that generate images and video, as well as audio and speech. And it can train all of them, using the same fast kernels that the original library was built around. But the app does more than just run and train models. It has a built-in web search and a deep research mode. It can also execute code in a sandbox and make tool calls. And it plugs directly into coding agents so you can point a local model at your project and swap models without changing your workflow. The API it exposes speaks the same language as OpenAI’s, so existing scripts and apps can connect to a local model without being rewritten. The development trajectory of Unsloth is as clear as it is ambitious. The library was a tool for engineers who work in code. Studio removed the code. And Desktop is a full workstation that goes beyond fine-tuning into running models, generating images and video, doing research, and talking to other tools. Each step has lowered the barrier and widened the audience. The company’s stated mission is to help builders create custom models faster and better, and the products are moving steadily toward making that possible for people who would never write a training script. What people are saying about it Initial reviews of the software are mostly positive. One write-up in Towards AI framed Unsloth’s contribution as genuinely democratizing model customization, and a developer community article put it on a list of essential open-source libraries to know. On Hugging Face, where the models are published, users are enthusiastic about the web search and the code execution capabilities. There are a few caveats. One user reported trying the desktop app and going back to LM Studio because it was missing too many advanced inference settings, which is a sentiment I can echo. And one reviewer made a critical technical point worth mentioning: the speed figures Unsloth publishes are training benchmarks from specific model tests, and not a guarantee of how fast the desktop app will actually feel on your hardware. A preview of how this will feel I honestly think Unsloth Desktop might be the first desktop app that shows how we will interact with AI in the future, and specifically how we will interact with local AI. It is the first time I have been able to complete an entire Substack post, from research and drafting through editing to image generation, effectively in a single application that runs on my own machine and sends nothing anywhere. (I am still using ElevenLabs for the voiceover, merely out of convenience.) In my testing, I encountered only a few minor issues. The deep research function timed out on me a few times. I have heard from other users that they were having the same experience, which leads me to believe that this is probably something that will be fixed in an update. And there is the obvious speed issue. Because everything runs locally, it will not feel as fast as the online foundational models you are used to. A home computer is not a data center, and you will experience the difference when you are waiting for a long response. One thing I found particularly exciting is that the app works well with MCP servers, the standard that lets AI tools talk to external resources. I was able to connect Qwen3.8 to Consensus, a service for searching and verifying academic sources, and it handled the search and the verification as expected. So where does that leave me? Well, I do not see myself switching completely to Unsloth Desktop right now. The beta has a few rough edges and the slower speed is a real downside. But the system has every feature I would ever need. In combination with an advanced model such as Qwen3.8 27B it does the research, writes, and edits exceptionally well. And it can also generate images, video, and audio with the help of open-source diffusion models. Everything runs locally, is open source and completely free. I want to encourage you, and I mean this especially for educators, to go download Unsloth Desktop and see what it does. If you have a reasonably powerful computer, you will be surprised by its capabilities. To be clear, this post is not sponsored and I have no relationship with the company. I just think this is an outstanding tool that demonstrates clearly what local AI already can do today. Open-source tools and models have improved significantly, and the best way to understand how far they have come is to run one on your own machine and test it yourself. This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.theaugmentededucator.com/subscribe

About

Stories From Education's AI Frontier. Exploring the promises, pitfalls, and possibilities of algorithmic teaching and learning. An AI-voiced companion to the Substack of the same name. www.theaugmentededucator.com