Cosmos-Predict2.5: world simulation for Physical AI

A useful open release of video world models aimed at robots and driving, with a compelling robot-augmentation result. Its primary evaluation is a benchmark the same company authored, the robotics experiment is underpowered, and nothing in the recipe is ablated.

Paper: arXiv:2511.00062, NVIDIA, October 2025 Role: reviewer February 2026

Summary

Cosmos-Predict2.5 is NVIDIA's second generation of open video world foundation models for Physical AI. It introduces two families: Predict2.5 at 2B and 14B, a flow-based unified text-, image- and video-to-world generator, and Transfer2.5 at 2B, a ControlNet-style sim-to-real and real-to-real translation model that is 3.5× smaller than its predecessor while outperforming it. The training ingredients are a 200M-clip curated dataset surviving a 4% filter from more than 6B raw clips, domain-specific fine-tuning, model merging, RL post-training, and timestep distillation to four-step inference. The T5 text encoder is replaced with Cosmos-Reason1, a physical-AI-specialized VLM, and EDM diffusion gives way to flow matching. Demonstrated applications include robot policy augmentation through sim-to-real translation, multi-view driving simulation, camera-controllable multi-view generation, synthetic VLA training data, and action-conditioned world generation.

Strengths

  1. The first comprehensive open platform for Physical-AI world models. Two sizes, twelve specialized variants and a permissive license lower the barrier for robotics and driving researchers.
  2. Efficiency at scale. The 2B model is competitive on PAI-Bench with a 27B mixture-of-experts baseline at 93% fewer parameters, and human voting puts the 14B model ahead of a 14B competitor.
  3. The robot augmentation result is compelling. Transfer2.5 augmentation gives 24 of 30 successes across ten test conditions against 5 of 30 for standard image augmentation and 1 of 30 with none, on realistic adversarial conditions of novel objects, lighting and backgrounds, with semantic control through text prompts.
  4. Sound ablations in several places: RL post-training, the merging strategy, action-conditioning injection, and long-video degradation with a new relative metric.
  5. A thoughtful data pipeline. Seven curation stages, multi-level captioning, semantic deduplication and content sharding across 26 categories, with task-aware captions for robotics data.

Weaknesses

  1. PAI-Bench is NVIDIA-authored. It is the primary quantitative evaluation and the sole basis for ranking against external models, with no independent benchmark alongside it. For a venue submission that is a serious methodological concern.
  2. The robot policy experiment is drastically underpowered. One task, 100 demonstrations, three trials per condition, one robot and one camera, compared against a naive augmentation baseline rather than the field's domain-randomization methods. Three trials per condition cannot support conclusions across ten scenarios.
  3. No ablation of the full recipe. Flow matching, the new encoder, RL post-training, merging and domain fine-tuning arrive together, so nothing isolates which ingredient matters.
  4. Physical plausibility is claimed but weakly evaluated. A physics dataset is curated and described, yet no physics-specific metric appears and the benchmark's physics sub-scores are not reported.
  5. The comparison set is narrow. Only two Wan versions, omitting several contemporaneous open models.
  6. Table 12 is hard to read, with Transfer2.5 quality scores far below Transfer1's and ambiguous direction markers.
  7. The action-conditioned evaluation is limited to 100 Bridge episodes at low resolution and 5.8-second clips, with no closed-loop policy evaluation using the model as a simulator.
  8. Deformable and contact-rich dynamics are absent, which is exactly where a world model would be most valuable.

What I would ask the authors

  • How was the PAI-Bench test set kept free of contamination with Cosmos training data?
  • Were the base, baseline and augmented policies retrained for each comparison, and how sensitive are three-trial results to the seed?
  • What does Cosmos-Reason1 contribute as the text encoder relative to any VLM of similar scale?
  • What dominates the 96% of clips that the 4% filter removes, and does that bias the pre-training distribution?

Recommendation

Weak accept. The open release, the augmentation demonstration and the Transfer2.5 efficiency gain are genuine contributions. Self-evaluation on a self-authored benchmark, an underpowered robotics experiment and the missing ablation keep it short of a strong accept without revision. In practice the community has voted: Transfer2.5-2B was the most downloaded model of the family within months, which says the sim-to-real translation use case is the one people find useful.

Method. Written for a reading group in February 2026, with download figures and community discussion checked at the time with the help of an AI assistant. The assessment is mine.