A systems paper whose strongest contributions are in data engineering and in being the first open foundation model for humanoid manipulation. Its dual-system framing is more marketing than architecture, and its own successors moved away from it within nine months.
GR00T N1 sits where three threads meet: vision-language-action models, diffusion-based action generation, and cross-embodiment robot learning. It argues that humanoid robots need a full-stack solution of hardware, models and data, and answers with a dual-system architecture, a VLM for reasoning and a diffusion transformer for actions, trained on a data pyramid of internet video, synthetic trajectories and real teleoperation. The model is open-weight at 2.2B parameters and reports 76.8% success on real GR-1 humanoid tasks against 46.4% for a Diffusion Policy baseline.
The line iterated fast. N1.5 (May 2025) moved to a frozen Eagle 2.5 VLM with a simpler adapter, added the FLARE objective that aligns to future latent representations and unlocks learning from human egocentric video, extended embodiment support to single arms and gripper humanoids, and reported real GR-1 language following rising from 46.6% to 93.3%; on the Unitree G1 a post-trained N1.5 reached 98.8% on seen objects and 84.2% on novel ones. N1.6 (December 2025) doubled the DiT depth, switched the VLM to a Cosmos-Reason variant with native resolution, removed the adapter and unfroze the top VLM layers instead, and predicted state-relative actions. Each version became less dual-system: the VLM and DiT increasingly share parameters and training signal, converging on the tighter coupling that the π0 family already used. All weights stay under NVIDIA's non-commercial license.
Elsewhere, DreamGen systematized the neural-trajectory idea into a pipeline; VLA-0 claimed to beat π0, π0.5 and GR00T N1 on LIBERO with actions as plain text and no architectural changes, which challenges the premise that a separate action head is needed; and the LeRobot ecosystem made the Unitree G1 a platform on which ACT, Diffusion Policy, π0, π0.5 and GR00T can be trained on the same data pipeline, positioning GR00T as one option among peers.
The novelty is in the combination and engineering rather than any component. The data pyramid is the most original contribution: organizing training data by scale and embodiment specificity, extracting latent actions from human video, and generating neural trajectories with video models, which multiplied 88 hours of data to 827. The latent action space unifying action-less video, simulation and real trajectories is creative, with concurrent work elsewhere. Being the first open VLA aimed at humanoids with dexterous hands filled a real gap. The dual-system framing, borrowed from Kahneman, is a VLM plus a DiT with cross-attention; the paper is forthright that the backbone, loss, DiT, chunking and conditioning are all established.
A well-executed systems paper, strongest in data engineering and in opening the humanoid foundation-model space. The rapid move to N1.5 and N1.6, each stepping away from the original design, suggests the paper captured a snapshot of a fast-moving internal program rather than a settled architecture, and the jump from 46.6% to 93.3% on real language following makes the N1 numbers look like an early prototype. Its most lasting effect may be indirect: open weights seeded an ecosystem that commoditized the model layer and pushed NVIDIA's advantage toward its simulation toolchain.
Method. Written for a reading group in February 2026 in the archeologist role, which traces a paper's ancestry and descendants rather than judging its experiments. Release dates and benchmark figures were checked against the model cards and papers at the time.