Skild AI says its S1 model can teach a robot an unfamiliar, long-duration task from a single video demonstration. In NVIDIA's September 10, 2026 post, the system is described as a robotics foundation model that uses the video as its prompt, without updating its weights or undergoing task-specific post-training.
One video becomes the prompt
The workflow is deliberately simple: an operator records the task they want and supplies that video to S1. The model interprets the demonstrated intent, objects and sequence, then converts them into robot actions. Because the behavior is supplied in context, Skild AI says the robot can often handle tasks missing from its pretraining without being retrained for each one. NVIDIA calls this approach in-context learning.
That changes the handoff between a person and a physical system. Instead of translating every new procedure into a separate training cycle, the described process starts with a demonstration. The claim concerns how S1 receives and uses task information; it does not guarantee flawless execution in every environment.
Results reported by Skild AI
NVIDIA says S1 can execute unknown tasks lasting up to 10 minutes. The examples in the post include repotting a plant, making pancakes, preparing pour-over coffee and assembling kits. In a repotting test, Skild AI went from recording the demonstration to autonomous execution on hardware in 11 minutes.
In multi-step tests involving new tasks, S1 completed about 66% of the steps, compared with 9% for a similar AI system. NVIDIA presents that as an improvement of more than seven times. The measurement is step success, so it describes performance across the tested sequences rather than a universal success rate for all robotic work.
Skild AI also estimates that a short video example can be as useful as roughly 380 practical training examples. The company says collecting those examples manually could take 50 to 100 hours. Both figures are company estimates, describing an expected data-collection benefit rather than a guaranteed saving for every deployment.
From demonstrations to commercial deployments
The announcement places the model inside an existing commercial operation. At launch, Skild AI said it had reached a $100 million annualized revenue run rate 10 months after its first commercial deployment and had more than 60 deployment partnerships across manufacturing, logistics, inspection, security and food preparation, among other applications. Those partnerships put the S1 announcement in the context of active industrial and service deployments.
In a separate manufacturing example, Skild AI, NVIDIA and Foxconn are deploying Skild Brain on two-arm manipulators for precision assembly of NVIDIA Blackwell systems. In the workflow shown by NVIDIA, the robot installs a busbar and a limiter, tightens 16 screws and adapts to disturbances during the multi-step task. It is a concrete example of the physical coordination the companies are targeting, but it is described as a Skild Brain deployment rather than a separate S1 test.
One condition will shape how much the system improves after deployment. Skild AI says experience from commercial deployments can feed its broader model when customer agreements permit it. Data-use permissions are therefore part of the operating model: the learning loop can extend beyond the initial demonstration, but only where those agreements allow.
S1's reported advance is the move from task-specific retraining toward video demonstrations as the robot's immediate instruction. The available evidence supports a meaningful development in physical AI through the reported workflow, selected task tests and commercial footprint, not a blanket claim that robots can learn any job from any video.
