A video recorded by an operator could be enough to teach a robot a task it has never performed before. That is the premise behind Skild AI's S1, a new robot foundation model designed for industrial and logistics environments where procedures, products, and workstation layouts change frequently. The goal is to ease one of traditional robotics' primary constraints: any significant variation in the task typically requires data collection, development, training, and validation of new software.
S1 takes a video demonstration as an instruction, extracts the sequence and the objects involved from the example, and then attempts to transfer that intent to the robot operating in the field. According to Skild, the system requires neither model weight updates nor a post-training phase dedicated to that specific task. The company frames this capability as in-context learning: the model uses the context provided at runtime rather than being re-optimized for every procedure.
The announcement comes as part of a technological collaboration with NVIDIA spanning the entire development cycle: synthetic data generation, training, simulation, and deployment in physical scenarios. NVIDIA presents S1 as a use case for its Physical AI platform, the suite of hardware and software tools through which the group aims to support the development of machines capable of perceiving and acting in the real world.
From demonstrated gesture to action sequence
This requirement is not trivial. An industrial robot can be highly reliable when operating in a stable workcell and repeating a defined maneuver, but it tends to lose flexibility if components, packaging, tools, or the order of steps change. In such cases, demonstrating the task rather than explicitly programming every movement offers a potentially faster route, provided the system can distinguish what matters in the demonstration from the scene's incidental details.
Skild claims that S1 can handle novel tasks lasting up to ten minutes and comprising dozens of manipulation steps. Examples shared include repotting plants, making pancakes, brewing pour-over coffee, and assembling kits. These demonstrations involve diverse objects and contexts, but they share a core challenge: elementary skills must be chained in the correct order while maintaining awareness of progress through the procedure.
In tests released by the company, S1 recorded around a 66% single-step success rate on novel multi-stage tasks, compared to 9% for a comparable AI system. These results were provided by Skild AI rather than an independent evaluation, so on their own they do not establish how the model performs across different factories or across a broad range of robots. However, they indicate which metric the company views as critical: not just completing a controlled demo, but sustaining execution when the procedure requires many consecutive steps.
In a repotting trial, the team reports moving from recording the demonstration to autonomous execution on hardware in eleven minutes. Skild also calculates that a short video clip can deliver value comparable to roughly 380 physically collected examples—work that could take a human between 50 and 100 hours. While these are company estimates, they illustrate the economic benefit this approach targets: compressing the time required to bring a new task from the test bench to the production line.
The field test with Foxconn and Blackwell systems
The project is not confined to household or laboratory examples. Skild, NVIDIA, and Foxconn are deploying the Skild Brain on dual-arm manipulators for precision assembly of NVIDIA Blackwell systems. In one of the workflows demonstrated, the robot positions a busbar and a limit block, tightens 16 screws, and must continue operating even when disruptions deviate from the expected setup.
A task of this type brings together requirements that are often addressed separately: movement precision, contact control with components, sequence memory, and the ability to recover from an error or displaced objects. The value of the experimentation, if extended beyond the demonstration, lies precisely in verifying whether a foundation model can offer adaptability without sacrificing the guarantees that a production environment requires regarding timing, quality, and safety.
Skild claims to have reached an annual revenue run rate of $100 million ten months after its first commercial deployment and to have secured over 60 deployment partnerships. The cited areas include manufacturing, logistics, inspection, security, and food preparation. The run rate figure is not equivalent to revenue actually generated in a fiscal year, but it signals that the company is trying to convert the general-purpose model into contracts and operational installations across multiple sectors.
Synthetic data, simulation, and data collected from installations
Behind the idea of a single video lies much broader training. To interpret a demonstration that does not match what was previously seen, a model must have acquired robust representations of objects, actions, contexts, and robot configurations. Skild uses NVIDIA infrastructure to train its shared model by combining simulations, human videos, teleoperation, and, when permitted by customer agreements, data from commercial installations.
NVIDIA Cosmos is used to expand and structure the information base: the group's open world foundation models can transform videos into organized descriptions, while Cosmos Curator is used to annotate, filter, and sort data at scale. For simulation, Skild also relies on Omniverse and Isaac Sim, with Isaac Lab cited by the company as a key part of its research path. Simulation makes it possible to expose the system to numerous variations before deploying it on a physical machine, reducing the time and risks associated with direct testing.
This pattern also highlights a practical limitation of the announcement. Learning “from a video” does not mean that just any clip is enough to achieve reliable behavior anywhere. Footage quality, the match between demonstrated tools and available tools, lighting conditions, object reachability, and robot specifications can all affect the outcome. Furthermore, the leap from being able to recover from a perturbation in a test to handling every possible exception in a plant is still a transition that must be proven with more extensive production data.
Then there remains the question of validation. Reducing retraining can accelerate the introduction of new procedures, but it does not eliminate the need to verify that a procedure taught on the fly complies with site rules and safety requirements. In settings where robots work near humans or on expensive components, autonomy will have to coexist with operational boundaries, supervision, and dedicated controls.
For NVIDIA, S1 offers a concrete showcase for its Physical AI strategy: selling accelerated computing, simulation tools, and models capable of connecting digital data with physical actions. For Skild AI, the challenge will be proving that the promise of rapid adaptation holds up outside curated cases and delivers measurable benefits for those who frequently reconfigure lines, products, or setups. The deployment with Foxconn will be one of the most closely watched testbeds, as it brings the model into a high-precision assembly operation directly tied to NVIDIA hardware itself.



