Figure Introduces Helix 2.5, Bringing LLM-Style Scaling to Robot Learning
Figure’s robots tidied rooms, folded towels, and made beds across 30 homes without training in those spaces.

Figure has announced Helix 2.5 (opens in a new tab), a robot AI system tested on household chores across 30 unfamiliar Bay Area homes. The robots used each home’s furniture without collecting training data there.
The change is taking a learned chore into a new setting. A different bed, room layout, or set of objects no longer has to mean starting the training process again.
Learning the Task, Not the House
Figure adapted one base model for three behaviors, tidying living rooms, folding towels, and making beds. Each behavior was trained elsewhere before evaluation in the new homes.
The company calls this zero-shot generalization. Here, that means unfamiliar homes and objects, not chores learned without training.
What the Tests Showed
Figure reports 56% complete-task success with Index pretraining, compared with 9% for an otherwise equivalent model trained without it. Partial completion did not count, and safety interventions were failures.
The results build on Index, Figure’s human-behavior dataset (opens in a new tab), which we covered at launch. That effort gathers recordings of everyday work to help train robots.
Figure also reports that doubling the human training data predictably improved robot-action prediction, with model size held fixed. Smaller training runs helped forecast the largest run’s result, echoing the scaling approach used in language models.
The measured improvement concerns action prediction, rather than completed chores. It points toward a way to estimate what more human training data could deliver before committing to a larger training run.



