RoboPapers

Chris Paxton and Michael Cho

Chris Paxton & Michael Cho geek out over robotic papers with paper authors. robopapers.substack.com

  1. 1d ago

    Ep#110: ENPIRE: Agentic Robot PolicySelf-Improvement in the Real World

    In many fields, automated research has become increasingly plausible with the improving capabilities of frontier language models, and the field robotics may be no exception. In ENPIRE, Wenli Xiao, Jia Xie, and Tonghe Zhang build a harness framework that captures physical feedback and allows for automated robotics research, minimizing human effort while training policies that achieve a 99% success rate on challenging, dexterous tasks like organizing a pin box or fastening a zip tie. Learn more in Episode 110 of RoboPapers, with Michael Cho, Chris Paxton, and Ruijie He! Abstract Achieving dexterous robotic manipulation in the real world heavily relies on human supervision and algorithm engineering, which becomes a central bottleneck in the pursuit of general physical intelligence. Although emerging coding agents can generate code to automate algorithm search, their successes remain largely confined in digital environments. We conjecture that the missing abstraction to automate robotics research is a repeatable feedback loop for real-world policy improvement: reset the scene, execute a policy, verify the outcome, and refine the next iteration. To bridge this gap, we introduce ENPIRE, a harness framework for coding agents that instantiates this physical feedback routine with four core modules: an Environment module (EN) for automatic reset and verification, a Policy Improvement module (PI) that launches policy refinement, a Rollout module (R) to evaluate policies with one or multiple physical robots operating in parallel, and an Evolution module (E) in which coding agents analyze logs, consult literature, improve training infrastructure and algorithm code to address failure modes. This closed-loop system transforms real-world manipulation learning into a controllable optimization procedure, minimizing human effort while allowing fair ablations across training recipe and agent variants. Powered by ENPIRE, frontier coding agents can autonomously train a policy to achieve a 99% success rate on challenging, dexterous manipulation tasks, such as organizing a pin box, fastening a zip tie, and tool use, a process that further accelerates when we dispatch an agent team on a robot fleet. Our results suggest a practical and scalable path toward deploying coding agents to autonomously advancing robotics in the physical world. Learn More Project Page: ArXiV: https://arxiv.org/abs/2606.19980 Transcript Michael Cho (00:06.088)Hello everyone, welcome to another episode of RoboPapers. Today we’re super happy to have Wenli, Jia Xie and Tonghe. They are three of the co-authors of this paper called ENPIRE: Agentic Robot Policy Self-Improvement in the Real World. Thank you so much for coming to the pod. Maybe as a start three of you can share a little bit about yourself first and then you guys can go through the paper. Wenli Xiao (00:30.821)Okay, I’ll get started. Sorry. So my name is Wenli. I’m a final-year PhD at CMU, advised by Professor Guanya Shi. Previously I did two-year internship at NVIDIA GEAR Lab with Jim Fan and Yuke Zhu. And later I joined Physical Intelligence as intern. Yeah, I did a bunch of VLA training, post-training, real-world RL and agentic stuff and now I’m mainly doing pretraining, studying the science of training and data. Yeah. Michael Cho (01:14.488)Yeah, I think just yeah. Jia Xie (01:16.463)Okay, sure. Yeah. Hi everyone. My name is Jia and I’m a master’s student at CMU, advised by Guanya. Of course we’re labmates. And I’ve interned at xDof as a research intern over the summer. And I’ve been working on some agent stuff and the extent stuff lately. So yeah, thanks for having us. Tonghe Zhang (01:43.739)Okay. Hi, my name is Tonghe. I’m also a master’s student at Professor Guanya Shi’s lab. I was doing internship at Amazon this summer and previously I’ve been working on reinforcement learning. And thank you guys for inviting us. Michael Cho (02:01.15)Alright, perfect. Great to have you guys. Yeah, Wenli, I know you have a slide, maybe you can go through with us. Wenli Xiao (02:08.228)Okay, yeah. So yeah, today the main topic is to introduce our paper, but as we all know, recently there’s a major breakthrough in AGI and that affects robotics, that is GPT-6 Astra. There are a lot of demos from Twitter that shows people apply this model to do different tasks. That related to robotics, like real-to-sim, sim-to-real, in-context learning, and code-as-policy. And we are very excited and we are kind of a little bit frustrating and also wondering how well robotics been solved. But we will leave the discussion to later section. And because even before the GPT-6 Astra Already witnessed a huge improvement of coding agent this year because we felt like coding agent is a biggest thing in AI this year because starting from February or last year’s December we see like a lot of a lot very good coding agent like Opus and Codex. They can do a lot of stuff like starting from front-end to Back-end and build a lot of great applications. And actually, people have been trying to use these models in robotics every time there’s a breakthrough in AI. And for this year there’s a CaP-X project, which we apply the and benchmark the state-of-the-art LLM model and VLM model to Try to solve and hill-climb on the robotic simulation benchmarks and it did pretty well in simulation. And we also tried to use coding agent to just solve the very challenging manipulation benchmark. This is a very classical task which would you need to manipulate a dot to push a T-shaped object until you perfectly matches the target. Wenli Xiao (04:31.492)And this was really challenging and people find a solution or people didn’t find a solution until diffusion policy. But now we just prompt the coding agent out of the box. This is actually Codex. We prompted to write a heuristic policy with no neural network training to achieve a hundred percent success rate. It just try and try and eventually it can like Write very good heuristic b policy. And this applies to different coding agents like Codex Claude Code and Kimi from Moonshot. This was very exciting to me. And actually during the past two years there are many impressive works about applying coding agent for robotics. So here I just like briefly divide them into three categories, like VLA plus harness, so you can achieve great performance that can be better than either VLA or code-as-policy itself. And there’s of course Code-as-Policy, which people apply coding agent to try and practice in simulation so they can learn some skills and later apply to the real-world. Real world robotics. And there’s also physical autoresearch, which is the ENPIRE work that we are gonna discuss today. So the major difference between the first two categories and the physical autoresearch is even all of them applied the method in the real-world. For the first two part of first two categories, they need tons of human effort. So if we summarize the procedure of the researchers do these projects, it’s like they apply coding agent to write code and do some practice in simulation and also with some guidance of human, it develops some impressive skills. And then the researchers, PhD students, deploy the policy in the real-world and watch the results. Jia Xie (06:54.974)Okay. Wenli Xiao (06:55.682)And maybe for the first trial it doesn’t work, or even the robot just do nonsense. And then the student just describes what happened and type in the terminal, and then the coding agent fix it and do another trial error and apply and suddenly it works, and then can submit paper. So in this process we observe that humans are a major bottleneck. Because we become the agent’s actuators and sensors. Yeah, because the observation was from like six months ago, now with GPT 6 extra, maybe we don’t need to be the sensors of the robot. GPT can like if you set up a webcam in a third-person view, maybe GPT can observe that the policy doesn’t work. But six months ago we didn’t have such good GPT and we have GPT 5.5 and we still wanna see how much it can do. So here is an example, we assemble the robot arm used some DaMiao motors because we are too lazy. We didn’t want to implement the firmware and we didn’t wanna implement the low-level control. So We just plug in the USB-CAN to the computer that has Claude Code installed and we prompt it to implement the impedance controller for the arm and try it by itself. And with great expectation and the Claude just like just think and write, think and write, and after five minutes it finished. And here’s the result. Michael Cho (08:50.424)Gosh. Ouch. Wenli Xiao (08:51.676)Yeah. And we were w glad we muted this video. Yeah, we were screaming about it. Yeah, and then we took another night to redo the 3D printing and assemble the arm again. Yeah. Chris Paxton (08:54.443)Did that break the robot? Is it? Yeah. Michael Cho (09:03.74)My Michael Cho (09:13.212)Well, you guys were very brave, man. RJ He (09:13.527)I think you’re particularly making Chris sad, given his background. Wenli Xiao (09:16.732)Yeah. So well it’s a little bit frustrating, but because autoresearch works so well in simulation that it works and your capacity to make a Twitter post. And there’s a lost also a lot of great applications of applying autoresearch in sim and digital world. So we started to think Michael Cho (09:20.366)Ouch, ouch. Wenli Xiao (09:43.857)What’s different between autoresearch in sim versus in the physical world? And we think that’s because in simulation or in sandbox you can do reset for infinitely many times. You have a verifiable reward, you can heal climate. And of course there’s a sandbox, you won’t break anything. But that may not be true because maybe with A S I it can jailbreak the sandbox. But for robotics, we don’t even have this harness or sandbox in the real-world. So that’s the motivation of the entire project. So for this project we wanna take a

    Ep#110: ENPIRE: Agentic Robot PolicySelf-Improvement in the Real World
  2. 3d ago

    Ep#109: Enigma - Connecting 100 Robots To The Internet

    What happens if you connect a hundred AI-powered robots to the internet and just let people give them any instructions they want? Enigma founder Jonathan Jacobi joins us to talk about his experience deploying a fleet of robots on the internet, keeping them operational, and watching people truly test the generality and robustness of frontier robotics models. Watch Episode 109 of RoboPapers with Chris Paxton and Ruijie He today! Learn More Original Post on X Robots.online website (no longer active) Enigma.inc Transcript Transcript is automatically generated by Riverside and edited by ChatGPT. Chris Paxton (00:05) Hey everyone, and welcome to another episode of RoboPapers. Today we’re super excited to have Jonathan Jacobi from Enigma here to tell us about their work in foundation models and about putting a hundred robots online. So, could you introduce yourself? Jonathan Jacobi (00:19) Yeah. Thank you, Chris. Thank you, RJ. It’s a pleasure to be on RoboPapers. Big fans. A bit about Enigma: we’re a company working on foundation models for robotics, as Chris said, as well as building innovative new interfaces for robotics. How should we interact with robots? How should the interface feel to users? We’re also interested in the back-and-forth between interface work and research on the models. About a month and a half ago, we put a hundred real robots online. You could actually interact with them over the internet. They were running our models through a variety of interfaces, and you could simply ask them to do different things. So we’re here to talk about the models, the interfaces, and the experience we’ve built. Chris Paxton (01:21) Awesome. So what did people do with all these robots that you put online? Jonathan Jacobi (01:25) The basic concept was that we wanted people to feel robotics today. We all see a lot of demos. We see a lot of cherry-picked demos—some of them are cherry-picked, some of them aren’t—and we wanted people to feel that something is here right now. One of the problems is that not enough people have robots in their houses. So how do we make enough people feel that robotics is happening now, while we still have this challenge of not having enough robots deployed in the world for the average person to use? We thought quite hard about how to make that possible. One of the best ideas we came up with was: maybe we can set up a huge warehouse with a lot of robots, put our models on those robots, and let anyone online interact with them. We also got a very cool domain: robots.online. We tried to create enough different setups that one robot could be in one scenario, another robot could be in another scenario, and people would have enough flexibility to do random things that they could really feel the AI work we’ve done. They could ask the robot to do things themselves, without the logistical challenge of putting a hundred humanoids all over the world. That’s why we used tabletop robots—arms that we customized and designed ourselves, building on public open-source designs—and let people interact with them freely. Some robots were artists. People could paint. People could do pantomime. People did very unexpected things. Someone asked one of the robots holding a paintbrush to pretend that the brush was a broom. Jonathan Jacobi (03:50) So it actually swept the floor with the paintbrush. We have some very cool examples of that that we can walk through later. The concept was that, at Enigma, we focus on generalization—text, video, audio—and we have very different interfaces and input methods for our models. We wanted people to be able to experience that generalization directly and see that it’s working today. Some of the robots were artists. Some were doing chemistry experiments. We had one that could defuse a bomb. We had one that could do lightsaber-esque fighting. Not actual lightsabers. Don’t sue me. But basically, we let people do different things with different robots online. RJ He (05:01) Okay, cool. Can you describe a little more of the interface that was set up? It sounds like each one is a single manipulator fixed to a tabletop, and then the interface is a regular chat-prompt type of thing, as if you were speaking to an AI agent. Is that right? Jonathan Jacobi (05:21) Yeah. That’s a good question. We had a static manipulator. One of them was based on the SO-101 design. We took the SO-101 and said, “We want this to look really polished.” So we brought in incredibly talented industrial designers and worked with them for about a month or two. They were working full-time on designing this new version. The robot you’re seeing on the screen is an SO-101 that we customized. The colors, the enclosure, everything was custom-designed. You can see that even the cables are hidden within the robot rather than hanging around it. These were static manipulators, and each table had a different type of interface. The artist robot had a chat interface similar to an AI agent. But we wanted to make it less boring than simply asking the robot to paint or perform pantomime, so we turned it into a game. You could give it an idea and it would paint something associated with that idea, and then you had to guess what it painted. For pantomime, you could tell it to pretend that an object was something else, and the robot would mime with it. So that was one chat-based interface. Then we had other interfaces. Let me walk through some of the things we built. This is our company page. You can see more than a hundred robots here. This is part of the actual warehouse. Those are spare robots and leftover components that we built. Here’s one of the chat interfaces where people were doing pantomime. You could do drawing, pantomime, and a few other things. And here’s another interesting interface. You could tap objects directly on the screen and then ask the robot to do something with them. You could tap the green bottle and say, “Pick this up,” then tap the orange flask and say, “Pour this into this.” Jonathan Jacobi (07:45) You could say, “Throw this here,” or whatever you wanted. We wanted to introduce a new kind of interactive interface—not just text, not just teleoperation, not just audio, but something new. Our vision is that if you extrapolate this ten or a hundred steps forward, the interface almost feels like a game. It feels like you’re playing the physical world. Think about something like StarCraft. Imagine having an interface where you can control a hundred robots. You can tap, type, drag and drop, and build agents on top of that to automate entire processes. That’s where we imagine the world going. Controlling robots becomes incredibly intuitive: you can type, tap, drag, drop, maybe talk to them. Another interface was the bomb-defusing robot. You were on a mission to defuse a bomb, and while you’re trying to do it, an operator is shouting at you: “You have to defuse it faster!” So you’re under pressure while talking to and controlling the robot. People did very unexpected things with this. Someone said, “Knock on the bomb twice at the bottom.” Someone else said, “Pour the orange flask over the pink button.” The pink button wasn’t standing upright—it was lying on the table. Chris Paxton (09:45) Can I ask—what simulation is that? That actually looks pretty good. The water physics there. Jonathan Jacobi (09:50) This is the physical world. This is not simulation. Chris Paxton (09:52) That’s not simulation? Okay. Well, that makes sense. RJ He (09:54) No way. Jonathan Jacobi (09:56) Yes. Every one of these is the physical world. None of this is simulation. Chris Paxton (10:01) Right. I knew the others were real. I don’t know why that one looked like a simulation to me. Sorry for doubting you there. RJ He (10:05) Yeah, agreed. Actually, maybe that’s a good question. First, what was the main motivation for doing this whole thing? Was it about the human-robot interface, or was it about seeing what kinds of prompts and questions people would ask? And then the second question is about this simulation-versus-real-world issue. Why did you need to do it in the real world, as opposed to simulation, which might have allowed you to scale even further? Jonathan Jacobi (10:34) I can start with the last question. You guys are at the heart of robotics. You see and feel all the ups and downs: “This is going to work. This is AGI. It’s almost there. It’s not quite there.” You feel that sine wave of robotics enthusiasm. We wanted to say: guys, we think it’s amazing that robotics is accelerating this fast, but we think it’s time to stop showing only video demos and let people actually interact with it. We wanted to reach millions of people. We ended up getting around two million views and millions of robot interactions. People were going back and forth with the robots millions of times. We didn’t want the demo to just be a video. We wanted it to be something interactive. That was one of the biggest motivations. We didn’t want to do it through simulation. We didn’t want to do it through a video. We wanted to put something out there that people could actually use today. We put a hundred real robots online. The setup you’re seeing here—the thing you thought was simulation—is the real booth. And we had hundreds of robots spread around this warehouse. This was a huge operation. We had to set up air conditioning, electricity, Wi-Fi—everything. But the real motivation was to have people actually interact with robots. The second part is what we wanted to learn from it. As a company, our thesis is that we’re research-first. Most of our team is doing research on the model side. But one thing I think is incredibly important—and that was true for language models, and I think is still lacking in video models—is that the interface really matters. If you look at language models, we had GPT before we had ChatGPT. Jonathan Jacobi (12:56)

    Ep#109: Enigma - Connecting 100 Robots To The Internet
  3. Oct 1

    Ep#108: VLS: Steering Pretrained Robot Policies via Vision–Language Models

    Pretrained vision-language-action models for robots often fail when a task is performed in a slightly different environment. With Vision-Language Steering, you can get better generalization performance out of a frozen robot policy, by steering the sampling process of your pretrained diffusion or flow-matching policy to handle out-of-distribution changes like obstacle position or changing visual appearance. Shuo Lin and Ishneet Sukhviner Singh join us to talk about how to better generalize robot policies with no new training data. Watch Episode 108 of RoboPapers today, with Jiafei Duan and Ruijie He! Abstract Why do pretrained diffusion or flow-matching policies fail when the same task is performed near an obstacle, on a shifted support surface, or amid mild clutter? Such failures rarely reflect missing motor skills; instead, they expose a limitation of imitation learning under train–test shifts, where action generation is tightly coupled to training-specific spatial configurations and task specifications. Retraining or fine-tuning to address these failures is costly and conceptually misaligned, as the required behaviors already exist but cannot be selectively adapted at test time. We propose Vision–Language Steering (VLS), a training-free framework for inference-time adaptation of frozen generative robot policies. VLS treats adaptation as an inference-time control problem, steering the sampling process of a pretrained diffusion or flow-matching policy in response to out-of-distribution observation–language inputs without modifying policy parameters. By leveraging vision–language models to synthesize trajectory-differentiable reward functions, VLS guides denoising toward action trajectories that satisfy test-time spatial and task requirements. Learn More Project Page: https://vision-language-steering.github.io/webpage/ ArXiV: https://arxiv.org/pdf/2602.03973 Transcript Transcripts are lightly edited by ChatGPT. RJ He (00:01) Okay. Hi everyone. Welcome to another episode of Robot Papers. Today we are super excited to have with us Shuo Liu and Ishneet, who will be sharing their paper on Vision-Language Steering: Steering Pre-Trained Robot Policies via Vision-Language Models. Shuo, Ishneet, do you all want to introduce yourselves? Shuo Liu (00:29) Yes. Hi everyone. My name is Shuo Liu. For now I’m a research engineer at Cortex Robot AI and previously I graduated from University of Washington, where I was also a student researcher at Allen Institute for AI, advised by Professor Ranji Krishna and Professor Jiafei Duan, focus on robot learning from develop foundation models to deploying them into real world. RJ He (01:01) Isn’t it? Ishneet (01:02) Yeah. Hi everyone, my name is Ishneet. I work with Professor Jiafei on this paper, VLS, with Shuo. I’m currently at Menlo Research, where I also work on Asimov, and I’ll be going to Oxford for my undergrad later this year. RJ He (01:17) Super exciting. Amazing that you’re already having publications at this young age. It’s awesome. Okay, without further ado, Shuo, why don’t you take us through the paper? Shuo Liu (01:26) Okay, sure, yeah. Let me share the screen. Okay, so just make sure it’s it’s good to say right. RJ He (01:38) Yeah, it’s good. Shuo Liu (01:39) Yeah. Hi everyone. Today I want to share our work vision language theory. This work is finished when I was a master student at the University of Washington and advised by Professor Ranjik Krishna and Professor Jiafei Duan. And so today I want to introduce motivation of this work and the method we solve this problem and the experiment results and also some future works maybe could be done from this direction. So the problem is from the observation of some small front twining of the Like models like π0 or π0 Po, we found that current generative robot policies are can’t are kind of tightly coupled of the training specific configurations, so which cause some failing to decouple the reusable motor skills from the normal task constraints. Like for example, here we have some cubes on the table and if you want to stack them up to a tower with specific numbers you may call the π0.5 pre-trained value a policy to stack some cubes like if you want to stack two cubes and using the termination signal like the pushing the red button on the table it can do very well because you Because for this specific task you can fine twin the base policy a bit, like stack two cubes and push the button or stack four cubes to push the button. And but what if at test time you ask the base policy to stack three cubes? It’s never it’s never contained in the fine-tuning data set, so the base policy Shuo Liu (03:33) Can learn how to stack two cubes or stack four cubes and push the button, but it will never know how to st how to stack three cubes and push the button. This is because the training configuration we combine the language instruction and the image observation with the task two cubes and four cubes together and the polic the base policy is never contained with task specific instruction stack three cubes. So the motor skills like moving the gripper to the cube and stack it on another one and going push the red button is already included in the base policy. But at test time if you given another out of distribution scenario or out of distribution language instruction, the base policy can always fail. So we think this is from the standard imitation learning reset. So for now if we have a base policy like take π0.5 or Momo Act 2 for example, the base policy can be trained on the behavior cloning target, which means given an expert data set which contains observation, action, and the language for each index and the optimization target is to maximize the likelihood of the base policy to be to predict the next action chunk based on current observation and the absurd language instruction. So the objective is stand static and also is very distribution dependent, which means after fun training your base policy can only capture the special and semantic correlations on the training manifold if you change the observation or language at test time the base policy may confuse. This observation is clear at our daily work because if you take one pre-trained VLA policy you just don’t want to like use it in the Shuo Liu (05:53) In distribution tasks, you may change the scenario, you may change the language to the VLA. So for example, sometimes we may face observation out of distribution signal which means we may made some visual or special shift and some semantic shift which means the change of language distribution. So at that time the world may drift away from the training distribution and how we can solve it. We for now we have some two solutions. The first one is we recollect a small dataset out of distribution scenario and run another training cycle. But we need to repeat this circle for every new deployment and make the base policy useful for this specific scenario. But this is kind of misaligned with the target of we train very large VLA policy. That’s because the target is to make the base policy general generalizable but and we can using the pre-train and post-train recipe to make the skills like pick and place or open close this kind of low level skills already inside the VLA policy but we are paying the cost to relear the existing skills at test time s simply because we are not selectively adapt to the base policy at test time. So in this work we want to train we want to try another way that is steering the base policy at inference time, which means we’re taking the we’re taking the ability of generative models like photomatching or diffusion models which are Shuo Liu (07:46) Broadly used in current VLAs like the action part action expert part. We can take the advan advantag advancement of their ability to denoise iteratively. So for example the diffusion model is denoising to the target distribution from Gaussian noise and using many diff denoising steps and the flow matching are learning denoising paths to from the Gaussian points to the target distribution directly. So we can use the this the ar feature of denoising and we can add some guided guidance signals into the denoising parts to change the direction of denoising. But the challenge is can we find the right way to steer the flow matching and the diffusion policies to K and well kept their in-distribution performance. The way we can use is using the classifier guidance which means we can inject one gradient signal that push the denoising trajectory towards the out of distribution con condition, which is from the feature of diffusion policy and flow matching policy. We’re we’re not going too deep to of the mice today. But they do have this feature to be guid to be guided in the densing process. So here I want to introduce the real challenge if we can have a guidance signal to push the denoising process of these generative models, then how to make this such a guidance signal what should it to be? So in this work we introduced three requirements of this guidance signal. So firstly since the out of distribution observation and language they are structured so we need to encode the Shuo Liu (09:59) Out of distribution special and the semantic constraints into one into one like Into one feature and out one just class label because you cannot you cannot classification every new deployment based on the functioning scenario. And so we need so we are requiring two needs for the guidance signal. So first one, it should be correctly interrupt the geometry and logical structure of the out of distribution condition and secondly it should provide one dense informative gradient to the action trajectory diffusion process. So based on this tunes we designed one method to provide one very good guidance signal to answer these two questions and before we go into RJ He (10:54) Yeah. So maybe I can just jump in here. What kind of OOD conditions are you mostly conce

    Ep#108: VLS: Steering Pretrained Robot Policies via Vision–Language Models
  4. Sep 25

    Ep#107: Do As I Do: Dexterous Manipulation Data from Everyday Human Videos

    Achieving human-like dexterity is the next frontier for robotics, and yet dexterity data is often subtly hard to scale. Real-world dexterity data, including things like finger-pose estimates, is often slightly off, making it physically invalid and hard to execute on real hardware and hard to learn from. DO AS I DO is an algorithm for reconstructing and retargeting monocular RGB videos to robot hands, outperforming the state of the art and even working from generated videos. Bhawna Paliwal, Haritheja Etukuru, Will Liang, and Mahi Shafiullah join us to tell us more. Watch Episode 107 of RoboPapers, with Chris Paxton and Jiafei Duan, now! Abstract How can we scalably generate data for robotic manipulation, especially on human-like platforms such as dexterous multi-fingered hands? Learning from human videos has recently emerged as a likely answer to this question. However, difficulties in estimating hand-object interaction and crossing the human-to-robot embodiment gap have hindered the adoption of abundant monocular RGB-only human videos as the primary source of robot manipulation data. In this work, we present DO AS I DO, an algorithm to reconstruct and retarget monocular RGB human videos to multi-fingered dexterous robotic hands. DO AS I DO reconstructs hand-object interactions from various egocentric and exocentric in-the-wild video sources. The algorithm then retargets these hand-object interaction estimates into a sequence of actions executable in the real world, yielding robot-complete manipulation data from disparate human videos. Overall, DO AS I DO outperforms previous state of the art in estimating hand-object interactions and extracting dexterous manipulation trajectories from RGB videos, as we show in experiments on datasets with ground truths and on a dataset of video clips collected online. Our experiments enable us to propose an efficacy playbook for practitioners collecting human data for manipulation. Learn More Project Page: https://do-as-i-do.com/ ArXiV: https://arxiv.org/abs/2606.19333 Github: https://github.com/malik-group/do-as-i-do Original thread on X Transcript Lightly edited for transcription errors, names, technical terminology, punctuation, and readability. Conversational phrasing is otherwise preserved. Editing done automatically with ChatGPT. Chris Paxton (00:05.479) Hey everyone, welcome to another episode of RoboPapers. Today we have some great work from Bhawna, Haritheja, Will, and Mahi: “Do as I Do: Dexterous Manipulation Data from Everyday Human Videos.” Welcome to the show, guys. Could you introduce yourselves quickly before we get started? Bhawna Paliwal (00:24.977) Yeah, hi, I’m Bhawna. I’m a PhD student here at UC Berkeley, advised by Jitendra Malik. Haritheja Etukuru (00:32.692) I’m Haritheja. I’m a second-year PhD student at Berkeley, also advised by Jitendra. William Liang (00:37.929) I’m Will. I’m also a second-year PhD student, advised by Jitendra and Pieter. Mahi Shafiullah (00:42.606) Hey everybody, I’m Mahi. I’m a postdoc currently at Berkeley, and I’m also spending some time at Amazon FAR right now. Chris Paxton (00:50.983) Awesome. Well, yeah, would you guys like to tell us a little bit about this? What have you guys been working on? Bhawna Paliwal (01:01.499) So I’ll just start with a little introduction to the paper. So mainly the motivation behind this work was how can we utilize the massive video data available on internet for training robots. And most of the prior work has been which which has involved learning from human videos, has involved co-training with robot data, which which eventually involves having like hundreds or thousands of hours of teleop data combined with, let’s say, twenty thousand hours of video. So our objective in this paper is to understand. Whether there is some intermediate representation, let’s say in in 3D, where we can extract the right kind of information from human videos and make learning from human videos much more efficient. So essentially every human video where hand-object interaction is going on turns into a separate trajectory on which you can train like a robot model. So prior work has looked into this problem mostly through by collecting their own data and maybe having like specific settings where they’re They have like specific angles and specific robot specific hands hand-object tasks available. But here we look at internet videos in general where there might be a lot of occlusion, motion blur and so on, which are not necessarily recorded for robot learning. Yeah. Chris Paxton (02:21.331) Right. Some of the videos here in the highlights look incredibly good, particularly with the dexterous hand. So how does this work? Bhawna Paliwal (02:35.113) So I think As like broadly there are like two steps. First we have reconstruction where we reconstruct the object in 3D, as shown here. And here and then we retarget it onto a robot hand, in our case the Sharpa hand. And for reconstruction, like the hardest part was object tracking. That is most of the times when we are looking into hand-object videos, the objects are mostly occluded, especially the tiny objects such as this one. There are like just fewer than a hundred pixels where the object is actually visible. And reconstructing the object becomes hard first. And secondly, tracking the object becomes even harder, even if you know the shape. So, in this work, we first we utilize like all the computer vision progress that has been made recently, especially models like SAM 3D, which can just like take a single image and give you the shape and pose of the object in that particular frame. And then we come up with our own method of how do you utilize the SAM 3D kind of methods to also track the object over time. And the motivation behind was just behind that was just this that SAM 3D has been trained on lots of images and artificially constructed images where like some particular object is blurred out or it has been hidden by some other object. And then so it knows when the how the object looks when it’s like heavily occluded. And then what we do is like we predict the shape and pose at one particular frame and then propagate it to next into like further frames. while keeping the constraint that the shape of the object should not change too much and things like that. So I let Will explain the retargeting part. Yeah. So Jiafei Duan (04:18.402) Yeah, I think I have a sorry. I have a question. Yeah, so I actually I read I read the paper quite a few times. and I find it very interesting. So what we are seeing here is an open-loop replay, is that my understanding? Right. So would that be like any limitation that you guys foresee? I mean, maybe I’m jumping ahead of time, but I’m just curious, right? Like for example, certain tasks when you are, you know, doing requires some level of precision or perception, then how will you address that? Or is it just more of the motion that matters at this point? Bhawna Paliwal (04:51.655) Yeah, so these trajectories are right now open loop and the how we foresee this should be changed into a perceptive task, which is definitely necessary, is by training taking this reference, domain randomizing the object position, and then training an RL policy, let’s say using point cloud as input. So this will involve just like since you already have the reference trajectory in simulation, you also know the object shape, and then you can like train an RL policy. The object pose is changing and let’s say like even the your environment is changing, lighting and so on, and then you can do all these kind of domain randomization tricks to kind of make it work for one object. And then once you have lots of such policies, you can probably distill it into one big model with let’s say behavior cloning. yeah. Mahi, maybe you would like to add more things. Mahi Shafiullah (05:46.155) Yeah, I mean I think one of the things that I think about when it comes to this project is really that we’re exploring how to get robot data from kind of arbitrary internet videos, right? But then like robot data is not the same as a robot policy. Imagine when you’re doing behavior cloning, you still have to do the same t thing like twenty, twenty-five, two hundred times. And that is because just a single trajectory does not really make a policy. It kind of like gives you what to do in a certain situation, but not what to do in small variations of the same situation. And you have to robustify around that trajectory somehow. multiple ways of doing this, right? Like the behavior cloning people likes to learn a map from observation to actions and the reinforcement learning people likes to train a policy that converts that from observation to action, maybe using some as reward, something as a reward. So What Rabna just described is more of an approach of like, if you wanna, you know, derive a reinforcement learning policy out of it, like a trajectory, how could he do it? But overall, I think the twenty thousand foot view that I personally take is that data is data and you want to figure out like how to make data useful. It could just be that we throw everything into a big policy where the images go in as an input and the actions, the robot actions that we’re deriving from this comes out as an output. And once you have enough scale, then that just learns the general policy for articulating a robot. So it’s pretty much open question, like how do you want to convert a robot trajectory into a robot policy? And I don’t want to like say that like a certain direct approach is better than the other. My personal bias at this point is what kind of Bhawna described, which is maybe we do some sort of sim-to-real. And especially for dexterous hands, I feel like sim-to-real is like a pretty good way to go about it. But You know, there are other papers which have kind of converted this into kind of like a two-finger gripper data and there you can actually do

    Ep#107: Do As I Do: Dexterous Manipulation Data from Everyday Human Videos
  5. Sep 23

    Ep#106: Flex-π: A Multi-Stream World-Action Model with Compute Flexibility

    World Action Models are becoming more popular in robotics, as they learn to predict the world jointly with learning how to act on it. However, these world predictions are usually purely based on reconstructing color images from video. This is limiting, because color is far from the most important quality for a robot moving around in the world — more important are qualities like 3D geometry and object semantics. In Flex-π, Ge Yan, Jesse Zhang, and team train a 6 billion parameter world action model to do exactly this, predicting 3D pointmaps and DINO features along with color. This results in a policy which is much more demonstration-efficient, generalizes well, and can perform complex long-horizon tasks. Learn more in Episode 106 of RoboPapers, hosted by Jiafei Duan and Ruijie He! Abstract World-action models (WAMs) predict the future to act better, but nearly all of them predict only RGB latents, trained purely for pixel reconstruction, with no explicit signal for the 3D geometry or object semantics manipulation needs. We find a surprising free lunch: the same frozen video-generation VAE that encodes RGB also encodes 3D pointmaps almost losslessly, with no pointmap-specific training at all. This lets us supervise Flex-π, a 6B-parameter WAM, on 3D geometry and object-centric DINO semantics alongside RGB, at no cost in new sensors, new pre-training, or inference latency. Every visual signal is projected into this shared latent space and denoised jointly with actions inside a Mixture-of-Transformers backbone; per-stream dropout with cross-modality forcing then lets a single trained checkpoint run on any subset of these streams, from a fast action-only mode to full joint generation. The result is a policy that is exceptionally demonstration-efficient and generalizes well, beating the strongest baselines by up to 2-7× on dexterous, precise, real-world bimanual manipulation tasks both in and out of distribution, all while running faster than π0.5. Our project website: this https URL Learn More Project Page: https://flex-pi.github.io/ ArXiV: https://flex-pi.github.io/ Thread on X Transcript Transcript edited and prepared with automated transcription and checked by GPT 6. There may still be errors. Jiafei Duan (00:04.231) Okay, sure, sure, sure. Yeah, hello everyone. Welcome to a new episode of RoboPapers. We passed our hundredth paper mark and this is, I think, 102 or 103. And and today we have like two of the returning speakers from the previous episodes, Jesse and Ge Yan, to come on and really share about their new work, Flex-π. So maybe do you guys want to Reintroduce yourself a bit more and then I think for some of the audience who never heard the previous podcast before. Yeah. Ge Yan (00:38.151) Yeah, yeah, hi Jiafei, and thanks for having us. So I’m right now like a second year PhD student who is going to enter my third year at the University of Washington. And my PhD advisor is Dieter Fox. And now I’m working also closely with TRI and Jesse. Yeah, Jesse, it’s yours. Jesse Zhang (01:01.599) Yeah, I’m a postdoc at UW and I’m advised by Dieter Fox, and Abhishek. And so I started this Ge and I worked on this project for the past how long has it been? Maybe almost a year, yeah, it’s been a while. And it’s changed quite a bit since the initial version was definitely not a WAM. so it’s been very interesting how this has turned out. Ge Yan (01:18.833) Yeah, it’s been a long time. Ge Yan (01:26.567) Yeah, great. Jiafei Duan (01:27.145) Yeah, it’s it’s been it’s been a quite a while. I mean, I remember I was still there when all the people collecting data. Ge Yan (01:35.464) Yeah. Jesse Zhang (01:35.623) Yeah, yeah. I mean, even like Phil presented on on ideas that led to this and they were very different than what it is now. Jiafei Duan (01:42.677) Yeah, yeah. Yeah. I mean I mean for audience like just context, we are from the same group, so I I know exactly what this is about. But yeah, it definitely changed. But yeah, we hope to hear from Ge exactly what Flex-π is all about. So yeah, Ge, you want to take over. Ge Yan (01:49.629) Yeah. Ge Yan (02:00.219) Okay, thank you. Okay, hello everyone. Today we’re going to introduce Flex-π, a multi-stream world-action model with compute flexibility. I think we framed it as a policy that can do one checkpoint, any observed inputs, any generated features, and the speed accuracy operating point you pick and deployment. And this work is done by me, Jinghao, Yuzhi, Lori, Minwen, Jesse, and Dieter at the University of Washington. So I think recently we’ve seen a lot of work on on WAMs. I think it starts from the trend of very good video generation model, and people start to leverage those video priors from those very large video models and use it for action prediction. And we’ve seen a lot of progress, including like Fast-WAM, LingBot-VA, DINO most recent DINOTube paper work and and a lot of others. And I think we know that the WAMs are great, but they usually only predict RGB latent and that is specifically used to for reconstruction. And the previous work like DINO or DreamZero, so they really shows us that the world modeling objective itself enables better action prediction. And I think it is quite intuitive right now, like with with so many results show its benefits. So I think in our work we kind of want to ask a question like why not predict additional and that is especially manipulation-centric visual signals. So we really want to leverage those other signals because we want to further push the world modeling objective strengths, and we want to do some very high precision and dexterous tasks, and we want to achieve better demo efficiency and generalization. So the right is a a multi-view. Ge Yan (04:02.367) RGB video shown in a YAM setup that a model do this Soft-Bag Zipping task. And so right now you can see it’s purely RGB right? And so for Flex-π we focus on like two additional signals. So for better Demo efficiency and the generalization. So one is DINO. So we know the DINO had provides very good object-centric semantics that can be very good for generalization, especially in those cloud cloud environments. And also the 3D point. I think this 3D information really provides a lot of demo efficiency, lots of model to achieve sometimes even have high performance with much less data. As as we know, there’s no like free lunch, right? as we introduce like additional signals, it can like possibly like always introduce like additional cost. So that that really comes to one problem, because the most WAMs the the inference mode is decided at training time. So like what kind of inputs the model takes and what kind of features the model needs to generate. And we know that predicting the future makes us better, but actually w when you generate those at the inference time make it pretty slow. And it’s usually we see a lot of model that joint and generate those features, it’s very hard to do real time. And especially with our additional visual signals, that makes even makes it even harder. So we we kind of like want to do like one model that can do both. they can do like action-only prediction, and then you can also predict the futures as other WAMs. And it can be decided at the deployment time. Ge Yan (06:00.741) And before introducing our exact architecture, we want to share that because we build on top of the Wan video model backbone. So this the bottom video shows after training on internet-scale videos, the Wan model, this backbone has a very good physical prior. You can see it can generate this looks reasonable like videos. Although there you can see there are still some artifacts, but it really sh it shows the model can learn like the r what waterwheel looks like a reasonable behavior, even though it’s not exactly the the right. And and one thing about the one backbone is like the VAE part. So VAE can be regarded as an encoder. So for a VLA you can have like any those image language encoder, and for these one models, we use something called a Wan 2.2 VAE. So it’s trained by you take a sequence of input videos and you have a many many convolutional networks and you compress that into latent space. So you get a more compact latent representation, and then you reconstruct RGB video from that. And that is exactly the prediction targets of most WAMs. So I think most people understand that the WAMs predict the pixels, but it’s not exactly accurate because they’re actually predicting the RGB latent. That is optimized for pixel reconstruction. And we can also call it like appearance. One thing we found out is this VAE is is actually largely pre-trained on RGB videos, but it can surprisingly transfer pretty well to pointmaps. So pointmap, you can imagine it as like the same shape as RGB, but the pointmap is for RGB that each each pixel like is RGB three channels, and for pointmap is just XYZ. So they have they share the same format. Ge Yan (08:08.372) And by normalizing, they can be within the same range. It’s just quite surprising that because this VAE can transfer it pretty well, as shown in this bottom right videos, that on the left is the ground truth, on the right is VAE reconstructed. And it’s pretty much losslessly. And this actually kind of like solves a big problem about leveraging that 3D information. Because a while ago, like when we do all the 3D learning, try to leverage those information. We all we’re always bothered by what kind of 3D encoder we should use because unlike 2D models, 3D encoding is still not mature enough. And this actually shows a very promising direction that it kind of allows to use those 2D pre-trained model applied to a 3D pointmap and it works pretty well. RJ He (09:04.802) So sorry, if you just go back. maybe the first question is is this therefore then limited to the one two point two model or does this observation apply to other backbones as well? Ge Yan (09:19.378) Yeah, that’s a gr

    Ep#106: Flex-π: A Multi-Stream World-Action Model with Compute Flexibility
  6. Sep 18

    Ep#105: Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-Planning

    Imitation learning, especially with interventions, has driven so much recent robotics progress. However, improving a policy via targeted interventions until it reaches a useful and deployable success rate is a time and labor intensive process. Instead, wouldn’t it be great if policies could improve on their own? That’s what Varun Giridhar and Animesh Garg join us to talk about. In Q-Planning, they start with a large policy like pi-0.5, and add a Q-function estimator to predict value instead of just actions, then use both successful and failed rollouts to update this Q-function online, then use it to guide sampling and trajectory selection. With just a few rollouts they can dramatically improve policy performance online. This provides a way to do really difficult tasks like inserting a credit card into a wallet, increasing success rate from 25% to 80% in just a few iterations. Learn more in Episode 105 of RoboPapers, hosted by Michael Cho, Chris Paxton, and Ruijie He! Abstract Behaviour Cloning (BC) has driven remarkable progress in robot manipulation, yet it is fundamentally limited by its inability to self-improve: a policy that fails cannot learn from that failure without additional human demonstrations. Reinforcement Learning fine-tuning offers a path to self-improvement but has proven difficult to scale to the multi-billion-parameter models underpinning modern robot policies. We propose Q-Planning, which equips a large visuomotor BC policy with a small off-policy Q-function. Because a Q-function estimates value rather than imitates actions, it can be trained on the same successful demonstrations as the BC policy and later absorb both successful and failed deployment rollouts, an asymmetry BC does not have. We exploit this asymmetry to enable value-guided action selection at inference (a single-step QQ-weighted average over BC draws) and online self-improvement that fine-tunes only the Q-function, leaving the BC weights untouched. On LIBERO and bimanual RoboTwin, ten iterations of self-improvement lift every benchmark score we tested (LIBERO-10 93 → 99%, RoboTwin 83.8 → 91.4%) and shorten successful episodes on the near-ceiling suites (LIBERO-Object, LIBERO-Goal). On two contact-rich bimanual real-robot tasks, the same loop (BC frozen, no human intervention) improves purely from its own deployment rollouts: stack-cups 40 → 90% and insert-wallet 25 → 80% in five iterations, whereas SFT on successful rollouts alone stalls at 55% and 30%. Under an identical online budget Q-Planning is the only method, among Best-of-N, filtered SFT, IBRL, DSRL, and DAWR, that improves stably from failures without training an auxiliary actor. Learn More Project page: https://q-planning.github.io/ ArXiV: https://arxiv.org/abs/2608.21204 This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit robopapers.substack.com

    Ep#105: Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-Planning
  7. Sep 14

    Ep#104: Light-Loco-Parkour: Versatile Perceptive Whole-Body Locomotion via Multi-Skill Distillation

    One of the key advantages of legged robots like humanoids should be how effectively they can move across a wide variety of terrain types to accomplish their task. But Light-Loco-Parkour from the team at Light Origins aims to change that: using only onboard sensing, they show a policy which can decide when to walk, vault, climb, or otherwise traverse as it moves through a complex environment. Unlike many others, it uses sparse seeds instead of relying on a large motion corpus, learning when to use its skills to move around without specific sub-task labels. Xiaodao Chen and Yuntao Ma join us to go into the details. Watch Episode 104 of RoboPapers now, with Michael Cho and Chris Paxton, to learn more! Abstract Existing humanoid whole-body control systems still fall short of the way humans move through cluttered terrain: they either track expressive whole-body references without terrain generalization, or react to terrain online while leaving the arms, torso, and knees largely unused. We present Light-Loco-Parkour (LightLP), an end-to-end perceptive whole-body locomotion system that closes this gap with a single deployable policy. Conditioned only on onboard depth and a velocity command, the policy decides when to walk, balance, climb, step down, or vault, with no reference input, skill label, hand-coded gate, or runtime motion graph. Compared with prior humanoid systems, LightLP makes three contributions. First, it introduces a whole-body perceptive-control pipeline that extends an RL-trained, velocity-tracking locomotion policy with parkour skills learned from object-interacting motions, so the same policy tracks velocity in open terrain, executes whole-body traversal at obstacles, and resumes locomotion afterward. Second, it acquires terrain-conditioned skills from sparse seeds by expanding a single motion into dynamically feasible, terrain-paired references across obstacle geometry, rather than relying on a large motion corpus. Third, it learns autonomous skill transitions from reward, letting the policy decide when and which whole-body skill to invoke from depth and command alone, with no one-hot skill label, hand-coded state machine, or runtime motion generator. Simulation and real-world experiments show high success across both benchmarked terrains and unseen obstacle variations, and the same policy transfers zero-shot to indoor and outdoor hardware experiments. These results demonstrate autonomous perceptive whole-body locomotion on a humanoid in outdoor settings, using only onboard sensing and a single deployable policy. Learn More Project page: https://light-loco-parkour.github.io/ Paper: https://light-loco-parkour.github.io/paper.pdf This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit robopapers.substack.com

    Ep#104: Light-Loco-Parkour: Versatile Perceptive Whole-Body Locomotion via Multi-Skill Distillation
  8. Sep 9

    Ep#103: Freeform Preference Learning for Robotic Manipulation

    Specifying reward functions for robots is one of the hardest things about reinforcement learning. Robot rewards often need to be very detailed; metrics like progress can be ill-defined and hard to estimate. This leads to most robot learning defaulting to sparse rewards or simple preference learning. But naive preference learning (having human annotators choose one trajectory over another) is an easy solution, but obscures a lot of the signal in complex tasks and can make learning a lot less efficient. Marcel Torné, Anubha Mahajan, and Abhijnya Bhat join us to talk about their solution: freeform preference learning, which lets annotators define natural-language axes to compare trajectories over. This improves real-world performance on long-horizon manipulation tasks over sparse rewards and simple binary preference learning. Watch Epsiode 103 of RoboPapers, with Michael Cho and Jiafei Duan, today to learn more! Abstract Reward design remains a central bottleneck for autonomous robot policy improvement, especially in long-horizon manipulation tasks where sparse success labels provide too little signal and binary preferences collapse many competing notions of quality into one ambiguous signal. We introduce Freeform Preference Learning (FPL), a method for learning robot policies from freeform human preferences. Rather than asking annotators which of two trajectories is better overall, FPL lets them define natural-language preference axes, such as speed, safety, quality of placement, or carefulness, and provide pairwise preferences along each axis. These annotations are used to learn a language-conditioned reward model that maps a trajectory and preference label to an axis-specific reward. We use this model to train a reward-conditioned policy that optimizes across the multiple human-specified dimensions. Across four real-world and two simulated long-horizon manipulation tasks, FPL improves over sparse-reward and binary-preference methods by 38 percentage points. Beyond improved performance, FPL learns dense progress signals without explicit subtask segmentation, shows compositionality of behavior not present in the data, and allows users to steer the policy towards different behaviors at test time without retraining. Blog post with videos available at this https URL Learn More Project page: https://freeform-pl.github.io/fpl.website/ ArXiV: https://arxiv.org/abs/2606.32027 This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit robopapers.substack.com

    Ep#103: Freeform Preference Learning for Robotic Manipulation

Ratings & Reviews

5
out of 5
3 Ratings

About

Chris Paxton & Michael Cho geek out over robotic papers with paper authors. robopapers.substack.com

You Might Also Like