Trending: On-device modelsSearch
iHeartGeek
iTECH

Figure's Helix 2.5 tidied 30 homes it had never seen

Figure's Helix 2.5 tidied, folded and made beds in 30 Bay Area homes with no fine-tuning, after Index pretraining lifted zero-shot success from 9% to 56%.

A white Figure humanoid with a glossy black head stands beside a person at a kitchen counter, one gloved hand resting on a grey towel next to a woven laundry basket.

Figure says Helix 2.5 can walk into a home it has never seen and start tidying. The humanoid robotics company evaluated the model in 30 Bay Area houses with no data collected in any of them, no fine-tuning and no adaptation to the rooms or the objects in them, and it credits Index, its human-behaviour pretraining dataset, with most of that ability.

The headline number is a jump in blind zero-shot success from 9% to 56%. Figure trained two policies on identical task data, one from random initialisation and one from the Index-pretrained Helix 2.5 model, holding architecture, optimisation, downstream data and evaluation fixed. Grading was whole-task only: every toy in the basket, every towel folded, the whole bed made, with no partial credit.

Three chores, 30 homes, one fixed checkpoint

The three behaviours are tidying a living room, folding towels and making a bed, and each was assessed against success criteria fixed before evaluation began. Living room tidy required 13 to 15 scattered toys to be picked up and placed in a basket. A single checkpoint was used across all 30 homes, no weights were adapted to the evaluation homes or objects, and Figure says no evaluation toy, towel or bedding appeared anywhere in the task-specification data, a check it ran with a model and then human review.

Half the data, same success rate

Figure also benchmarks Helix 2.5 against a Helix 02 policy trained for the same task in the environment where it was evaluated. Helix 2.5 matched that policy's success rate while using half as much adaptation data, and did it without ever having been in the rooms. The company highlights self-correction as another gain: stepping back to reposition, changing stance or walking around a bed to fix a fold, which matters more the longer a task runs.

A scaling law for human-to-robot transfer

The most consequential claim is that human-to-robot transfer scales. Figure trained four models on nested subsets of Index spanning an eightfold increase in pretraining data, with model size and downstream training held fixed, and says loss fell predictably with every doubling. The relationship was tight enough that the company forecast its largest run's test loss to four decimal places before training began, with a forecasting error of 0.54% of the variation across the full range. Figure calls it the first human-to-robot transfer scaling law measured on a humanoid.

What is not being claimed

Every figure here is Figure's own. The evaluations are internal, the demonstration videos are the company's, and no third party has replicated the result. Figure is explicit that the point is not that general humanoid robotics is solved, and the fine print matters: zero-shot refers to the homes and objects, not to the tasks, which are still specified through fine-tuning data collected somewhere else. Index, meanwhile, is now generating new human experience at a rate of roughly 35 minutes a second, and Figure says it has committed $3.5bn of compute to training Helix, which is the resource the whole argument rests on.

Our opinion

Robotics has spent years collecting its own data one warehouse and one demonstration rig at a time, and Figure's wager is that the industry was short of the wrong resource. If a humanoid can learn how homes work from human behaviour and then be handed a single new chore, the cost of teaching a robot anything collapses from a building full of teleoperators to a fine-tuning run. The scaling law is the part that would change planning rather than demos, because it offers something robotics has never had: a way to estimate what the next doubling of human data buys before paying for it.

The caveats deserve the same attention as the result. A 56% success rate on whole-task grading is a research milestone, not a product, and the company is both the author and the examiner of its own benchmark. Until someone outside Figure runs the same evaluation in homes the company has never touched, this is a strong internal signal rather than an established law. What is no longer in doubt is the direction: the robots are being taught by people, not merely by other robots, and that is a much larger addressable dataset.

What we know
  • Figure evaluated Helix 2.5 in 30 Bay Area homes with no data collected in any of them and no fine-tuning
  • Index pretraining lifted blind zero-shot success from 9% to 56% on whole-task grading with no partial credit
  • The model matched a Helix 02 policy for the same task using half as much task-specific data