Team Comet released code, two checkpoints and the rejection-sampling dataset behind their second place in the 2025 BEHAVIOR Challenge. I priced each claim, ran the two I could afford, and found that the free one was the most informative.
Comet took second place at NeurIPS 2025 with a Q-score of 0.2514 on the held-out test set, then reported 0.345 on public validation after the challenge, up from 0.192 after pre-training. The recipe is a π0.5 backbone, about 1.5K hours of trajectories, and three rounds of rejection-sampling fine-tuning (RFT). What makes the report worth an empiricist's time is the last line of its abstract: everything was released. Most competition reports are unreproducible by construction. This one shipped its artifacts, so the question becomes which claims survive contact with them.
Reproduction is not one thing. The report makes five testable claims, and their costs span four orders of magnitude.
| Claim | Their cost | My cost | Reachable |
|---|---|---|---|
| Pre-training (pt50) | 50k steps on about 40 H200s | No. The config is literally named _gpu40 | |
| RFT post-training | 3 rounds of about 8,500 rollouts | 3,500+ GPU-hours | No |
| Per-task evaluation of the released checkpoint | 30 to 60 GPU-hours | Yes | |
| pt12 versus pt50 | 30 to 60 GPU-hours | Yes | |
| Table 3 inference ablations | 10 to 20 GPU-hours | Yes, cheapest and sharpest | |
| Characterising the released RFT dataset | none | Yes, free |
Two things fall out of the table before any GPU is touched. The inference-time ablations are the highest-value target, because they are runtime flags and the released checkpoints make them testable with no training. And the released RFT dataset lets me interrogate the post-training claim, the one I categorically cannot re-run, for free.
BEHAVIOR-1K runs on OmniGibson, which runs on Isaac Sim, which draws the scene with hardware ray tracing. NVIDIA lists the A100 and H100 as unsupported; they have no RT Cores. The V100 is older still. This inverts the usual ranking of hardware: the report trained on H200s, but you cannot evaluate on one.
| Architecture | Example | RT Cores | Runs the benchmark |
|---|---|---|---|
| Volta | V100 | No | No |
| Ampere GA100 | A100 | No | No |
| Hopper | H100, H200, GH200 | No | No |
| Ampere GA102 | A40, A10, RTX 3090 | Yes | Yes |
| Ada | L40S, RTX 4090 | Yes | Yes |
| Blackwell | RTX 5090 | Yes | Yes, on driver 580 |
My two best-specced machines were the two that could not do the job: a DGX with eight V100s and an A100 vGPU allocation. The A40 nodes on NCSA Delta were the only clearly RT-capable GPUs in the ACCESS-CI catalog. The national GPU buildout optimised for tensor-core throughput, and embodied-AI simulation fell outside it.
This explains a design decision in the paper. Section 4.3 rejects online RL partly because it needs GPUs with RT Cores for simulation and GPUs with Tensor Cores for training at the same time. RFT is not only a sample-efficiency choice. It lets training and simulation be decoupled in time, so you never need both GPU types at once. It is arguably the most transferable idea in the report, and it is buried in one sentence.
The report attributes the move from 0.224 to 0.345 to a refined task-balancing strategy when selecting 1,469 trajectories from about 25,500 rollouts. If balancing drove the gain, the kept set should be more uniform across tasks than raw rejection sampling would produce. I pulled meta/episodes.jsonl and meta/info.json from delinqu/comet-1.5k and computed per-task counts, per-task mean horizon, and the rank correlation between them. Zero GPU-hours.
The report credits task balancing. The artifact shows that balancing did comparatively little. What the data shows instead is a horizon-dependent sampling bias: rejection sampling harvests almost entirely from the short half of the benchmark and is silent exactly where the benchmark is hardest. That sharpens the report's own admission that sampling efficiency is too low. RFT cannot close the gap to the 0.611 theoretical best by construction, because it generates no signal on the eleven tasks that need it most.
Finding 5 also recasts the resolution result in Table 3. The gain from 0.30 to 0.60 is not from adding information; it is from ceasing to discard it. "High-fidelity perception is essential" is more honestly stated as "do not downsample your data."
Figure 4 of the report gives a success rate of 1.00 on turning_on_radio, the task the authors use for every cell of Tables 3 and 4. It is also the shortest task in the benchmark, 934 frames or 31 seconds, and a rate of 1.00 is easy to disprove. I evaluated the released pi05-b1kpt50-cs32 checkpoint with receding-horizon control and native resolution on all 20 public test instances, on two machines with different GPUs, drivers and installs.
| Machine | GPU | Report | n | Ours | 95% CI |
|---|---|---|---|---|---|
| marklxxxv | RTX 3090 | 1.00 | 20 | 0.00 | [0.00, 0.16] |
| b3iq | RTX 5090 | 1.00 | 20 | 0.00 | [0.00, 0.16] |
The policy failed all 40 instances, each running to the 4,300-step limit with a partial-credit Q-score of 0.000. A broken harness also gives zero, so the first question is whether the harness works. On training instances, which the policy saw during training, it succeeded on 1 of 4 with a mean Q of 0.25. The harness can record a success. The policy succeeds about a quarter of the time on data it trained on.
My first run was wrong, and the error was mine. I had changed the wrapper to 224×224 to match Table 3, and the released checkpoint scores zero at that resolution even on training data. The 224×224 numbers in Table 3 come from single-task models that are not public. The positive control caught this. Without it I would have reported a false result about the authors' work.
Three readings of 0 of 20 remain: the claim does not reproduce; the released checkpoint is not the model in Figure 4; or my setup is wrong in a way I did not find. The public record cannot settle the second. The checkpoint's build date and model card point to the 0.345 model, while its directory name and the README's model-zoo text point to the pre-training model, and the README calls the same files both things. The third is bounded by the control: a dead harness cannot succeed on training instances.
og.launch(), and a setup script that downloads the 2026 task instances while the evaluator reads the 2025 ones. Fixes for all three are in my fork.info.json frame count is 18% below the sum of episode lengths, probably not rebuilt after the last filter step. Nothing above depends on it.Method. The cost table, the dataset analysis and the evaluation runs are mine, carried out in September 2026. Fixes to the evaluation stack are in the fork linked above.