Introducing MilliVid, our new method for long-context video generation! MilliVid creates videos that are consistent over long time spans, without using retrieval heuristics or 3D maps! (1/n)
davidcharatan.com/millivid/#
Does scaling pre-training on general web video improve a complex manipulation task in real deployment?
We scale model size and pre-training compute, and test on one industrial task.
Yes. The better a pre-trained model predicts web video, the better its post-trained policy. 🧵
Based on his pioneering work on dataset distillation, my student George has noticed that his most recent method (to be released soon, stay tuned!), besides being SOTA in dataset distillation, also creates stunning synthetic "composites" of an artist's body of work. Check it out!
We discovered that our latest Dataset Distillation project can be used to create some beautiful synthetic images based on an artist's body of work!
Come see us at the @eccvconf Art Gallery starting today!
Explanation and some of my favorites in thread below:
1/ (Claude Monet)
At 10:20 am Malmö time, I will be speaking at the "X-Reason" workshop in the Palisades South room (turn right at registration, then down the stairs) about whether intermediate representations are important for embodied intelligence! #ECCV2026
I agree with Phil's take here: The progress of LLMs on controlling robots is quite interesting. Intuitively, this makes sense: controlling a robot is not so different from computer use, and an agent that is good at computer use is probably also good at controlling a robot and
Recently, there have been a lot of impressive demos of AI agents, like Claude, controlling robots.
I wrote a short blog post with my thoughts on the advent of these "robot-use agents."
web.mit.edu/phillipi/www/w…
I think it's an important change in the trajectory of robotics!