Skild AI’s S1 Teaches Robots New Tasks From One Video
Skild says the foundation model can follow unseen, ten-minute jobs without fine-tuning and reached 66 percent per-step success on its internal benchmark.

Skild AI has introduced S1 (opens in a new tab), a robotic foundation model the company says can learn a new physical task from one video and attempt it without any further training.
A foundation model is one AI system trained across many kinds of work. For S1, the new part is in-context learning. Instead of changing the model for each job, a person shows it what to do through a video prompt.
Show it once
Skild demonstrated S1 repotting a plant, cooking a pancake, making pour-over coffee, and assembling a kit. The company says none of those complete jobs appeared in its training data. Some ran for up to ten minutes and required dozens of steps.
The demonstration does not need to match the robot's exact view or surroundings. Skild says S1 works out the goal, connects objects in the video to objects in front of the robot, and tracks which parts of the job are finished.
For the plant task, only 11 minutes passed between the start of the human recording and S1's autonomous run. Most of that time went into capturing the example.
The robot is not meant to copy every motion. In Skild's examples, S1 used a cup when the watering can from the prompt was missing, added only a little juice when a glass was nearly full, and avoided an egg-dropping mistake made by the demonstrator.
The benchmark favors video prompts
Skild also compared S1's video prompting with a model given language instructions. The company says both policies used the same architecture, computing budget, and training data apart from the way each prompt was encoded.
On unseen tasks after 100,000 hours of training data, the video-prompted policy reached a 66 percent average success rate per step. The language-prompted model reached 9 percent. Those are not full-task completion rates. Skild used human intervention to recover from failures so every step could be scored.
In a separate internal comparison, Skild estimates that one video prompt matched the performance of about 380 task-specific training examples. More training eventually won. A conventionally post-trained policy reached 86 percent after 2,000 demonstrations, compared with S1's 66 percent from one prompt.
The results also show that video does not always start ahead. With 1,000 hours of pre-training on familiar tasks, the language-prompted model scored 53 percent and the in-context model scored 43 percent. Skild says the video approach improved faster as the data grew.
The next proof is outside Skild's lab
The demonstrations and benchmark come from Skild. The company has not released S1 for independent testing, published the model's full architecture, or opened its benchmark data. It says later reports will explain how S1 is trained.
Skild says S1 is already working with commercial partners and has opened a form for early access and deployment questions. It has not named the customers using S1 or announced general availability or pricing.
If outside tests reproduce the same jump from one demonstration to useful robot behavior, teaching a machine could begin to feel less like programming and more like showing a coworker the job. Independent trials across different robots and workplaces are the next evidence to watch.



