Good question thanks Ani. Here's a video that gives a sense of the answer here about the effect of prompting:
- in the first part, the model is just babbling with no prompt in context
- at the end, we add a prompt for taking money out of a pouch (generalizing to the wallet)
Very cool work, Pete! I'm curious if you've tried an extreme version of this: no in-context example at all; just place objects in front of the robot and see if it can infer the task.
I suspect this will have non-trivial success rates, which would also allow you to figure out
Last week @willknight came by and got a preview of one-shot prompting and we had some fun improvisational moments with the robot. Thanks Will for helping capture the moment.
I'm usually wary about talk of "the ChatGPT moment for robots" because language is less complex than the physical world. After visiting @GeneralistAI, though, where I saw their robots perform impressive manipulation with just a prompt, I think you can see it happening in a
Ever since I started working on robot foundation models a handful of years ago, the broad ability to one-shot in-context learn has been the single most vivid goal in my mind along the long road ahead. It’s something @andyzengineer and I have especially been thinking together on
Introducing GEN-1.5, a one-shot learner.
It can learn new tasks in a few seconds. Show it what to do, and it generalizes.
This capability emerged from pretraining on physical data at scale, as a step towards our mission of building general intelligence for the physical world.
We've improved how GEN-1 learns to adapt to new actuators and new robots at the lowest level, with up to 10-20x gains on internal benchmarks. This significantly boosts performance on high-precision tasks like disassembling parts from a NIST board.
Read more about GEN-1 in our