World Action Models are becoming more popular in robotics, as they learn to predict the world jointly with learning how to act on it. However, these world predictions are usually purely based on reconstructing color images from video. This is limiting, because color is far from the most important quality for a robot moving around in the world — more important are qualities like 3D geometry and object semantics. In Flex-π, Ge Yan, Jesse Zhang, and team train a 6 billion parameter world action model to do exactly this, predicting 3D pointmaps and DINO features along with color. This results in a policy which is much more demonstration-efficient, generalizes well, and can perform complex long-horizon tasks. Learn more in Episode 106 of RoboPapers, hosted by Jiafei Duan and Ruijie He! Abstract World-action models (WAMs) predict the future to act better, but nearly all of them predict only RGB latents, trained purely for pixel reconstruction, with no explicit signal for the 3D geometry or object semantics manipulation needs. We find a surprising free lunch: the same frozen video-generation VAE that encodes RGB also encodes 3D pointmaps almost losslessly, with no pointmap-specific training at all. This lets us supervise Flex-π, a 6B-parameter WAM, on 3D geometry and object-centric DINO semantics alongside RGB, at no cost in new sensors, new pre-training, or inference latency. Every visual signal is projected into this shared latent space and denoised jointly with actions inside a Mixture-of-Transformers backbone; per-stream dropout with cross-modality forcing then lets a single trained checkpoint run on any subset of these streams, from a fast action-only mode to full joint generation. The result is a policy that is exceptionally demonstration-efficient and generalizes well, beating the strongest baselines by up to 2-7× on dexterous, precise, real-world bimanual manipulation tasks both in and out of distribution, all while running faster than π0.5. Our project website: this https URL Learn More Project Page: https://flex-pi.github.io/ ArXiV: https://flex-pi.github.io/ Thread on X Transcript Transcript edited and prepared with automated transcription and checked by GPT 6. There may still be errors. Jiafei Duan (00:04.231) Okay, sure, sure, sure. Yeah, hello everyone. Welcome to a new episode of RoboPapers. We passed our hundredth paper mark and this is, I think, 102 or 103. And and today we have like two of the returning speakers from the previous episodes, Jesse and Ge Yan, to come on and really share about their new work, Flex-π. So maybe do you guys want to Reintroduce yourself a bit more and then I think for some of the audience who never heard the previous podcast before. Yeah. Ge Yan (00:38.151) Yeah, yeah, hi Jiafei, and thanks for having us. So I’m right now like a second year PhD student who is going to enter my third year at the University of Washington. And my PhD advisor is Dieter Fox. And now I’m working also closely with TRI and Jesse. Yeah, Jesse, it’s yours. Jesse Zhang (01:01.599) Yeah, I’m a postdoc at UW and I’m advised by Dieter Fox, and Abhishek. And so I started this Ge and I worked on this project for the past how long has it been? Maybe almost a year, yeah, it’s been a while. And it’s changed quite a bit since the initial version was definitely not a WAM. so it’s been very interesting how this has turned out. Ge Yan (01:18.833) Yeah, it’s been a long time. Ge Yan (01:26.567) Yeah, great. Jiafei Duan (01:27.145) Yeah, it’s it’s been it’s been a quite a while. I mean, I remember I was still there when all the people collecting data. Ge Yan (01:35.464) Yeah. Jesse Zhang (01:35.623) Yeah, yeah. I mean, even like Phil presented on on ideas that led to this and they were very different than what it is now. Jiafei Duan (01:42.677) Yeah, yeah. Yeah. I mean I mean for audience like just context, we are from the same group, so I I know exactly what this is about. But yeah, it definitely changed. But yeah, we hope to hear from Ge exactly what Flex-π is all about. So yeah, Ge, you want to take over. Ge Yan (01:49.629) Yeah. Ge Yan (02:00.219) Okay, thank you. Okay, hello everyone. Today we’re going to introduce Flex-π, a multi-stream world-action model with compute flexibility. I think we framed it as a policy that can do one checkpoint, any observed inputs, any generated features, and the speed accuracy operating point you pick and deployment. And this work is done by me, Jinghao, Yuzhi, Lori, Minwen, Jesse, and Dieter at the University of Washington. So I think recently we’ve seen a lot of work on on WAMs. I think it starts from the trend of very good video generation model, and people start to leverage those video priors from those very large video models and use it for action prediction. And we’ve seen a lot of progress, including like Fast-WAM, LingBot-VA, DINO most recent DINOTube paper work and and a lot of others. And I think we know that the WAMs are great, but they usually only predict RGB latent and that is specifically used to for reconstruction. And the previous work like DINO or DreamZero, so they really shows us that the world modeling objective itself enables better action prediction. And I think it is quite intuitive right now, like with with so many results show its benefits. So I think in our work we kind of want to ask a question like why not predict additional and that is especially manipulation-centric visual signals. So we really want to leverage those other signals because we want to further push the world modeling objective strengths, and we want to do some very high precision and dexterous tasks, and we want to achieve better demo efficiency and generalization. So the right is a a multi-view. Ge Yan (04:02.367) RGB video shown in a YAM setup that a model do this Soft-Bag Zipping task. And so right now you can see it’s purely RGB right? And so for Flex-π we focus on like two additional signals. So for better Demo efficiency and the generalization. So one is DINO. So we know the DINO had provides very good object-centric semantics that can be very good for generalization, especially in those cloud cloud environments. And also the 3D point. I think this 3D information really provides a lot of demo efficiency, lots of model to achieve sometimes even have high performance with much less data. As as we know, there’s no like free lunch, right? as we introduce like additional signals, it can like possibly like always introduce like additional cost. So that that really comes to one problem, because the most WAMs the the inference mode is decided at training time. So like what kind of inputs the model takes and what kind of features the model needs to generate. And we know that predicting the future makes us better, but actually w when you generate those at the inference time make it pretty slow. And it’s usually we see a lot of model that joint and generate those features, it’s very hard to do real time. And especially with our additional visual signals, that makes even makes it even harder. So we we kind of like want to do like one model that can do both. they can do like action-only prediction, and then you can also predict the futures as other WAMs. And it can be decided at the deployment time. Ge Yan (06:00.741) And before introducing our exact architecture, we want to share that because we build on top of the Wan video model backbone. So this the bottom video shows after training on internet-scale videos, the Wan model, this backbone has a very good physical prior. You can see it can generate this looks reasonable like videos. Although there you can see there are still some artifacts, but it really sh it shows the model can learn like the r what waterwheel looks like a reasonable behavior, even though it’s not exactly the the right. And and one thing about the one backbone is like the VAE part. So VAE can be regarded as an encoder. So for a VLA you can have like any those image language encoder, and for these one models, we use something called a Wan 2.2 VAE. So it’s trained by you take a sequence of input videos and you have a many many convolutional networks and you compress that into latent space. So you get a more compact latent representation, and then you reconstruct RGB video from that. And that is exactly the prediction targets of most WAMs. So I think most people understand that the WAMs predict the pixels, but it’s not exactly accurate because they’re actually predicting the RGB latent. That is optimized for pixel reconstruction. And we can also call it like appearance. One thing we found out is this VAE is is actually largely pre-trained on RGB videos, but it can surprisingly transfer pretty well to pointmaps. So pointmap, you can imagine it as like the same shape as RGB, but the pointmap is for RGB that each each pixel like is RGB three channels, and for pointmap is just XYZ. So they have they share the same format. Ge Yan (08:08.372) And by normalizing, they can be within the same range. It’s just quite surprising that because this VAE can transfer it pretty well, as shown in this bottom right videos, that on the left is the ground truth, on the right is VAE reconstructed. And it’s pretty much losslessly. And this actually kind of like solves a big problem about leveraging that 3D information. Because a while ago, like when we do all the 3D learning, try to leverage those information. We all we’re always bothered by what kind of 3D encoder we should use because unlike 2D models, 3D encoding is still not mature enough. And this actually shows a very promising direction that it kind of allows to use those 2D pre-trained model applied to a 3D pointmap and it works pretty well. RJ He (09:04.802) So sorry, if you just go back. maybe the first question is is this therefore then limited to the one two point two model or does this observation apply to other backbones as well? Ge Yan (09:19.378) Yeah, that’s a gr