Paper review

Reviews and small experiments on papers in robot learning, written for a reading group. Each page ends with a note on how it was produced.

  1. Sep 2026
    Team Comet, second place in the 2025 BEHAVIOR Challenge · arXiv:2512.10071 · empiricist
    I priced five claims and ran the two I could afford. The released RFT dataset shows a cap rather than a rebalance and a strong horizon bias, and the released checkpoint scored 0 of 40 on the task the report puts at 1.00.
  2. Sep 2026
    Larchenko, Zarin and Karnatak · arXiv:2512.06951 · reviewer
    Honest and reproducible, but correlated-noise flow matching is asserted rather than ablated. Weak reject for a main track, accept as a competition report.
  3. Sep 2026
    Research note with preliminary probe results · empiricist
    72% of BEHAVIOR goal predicates are spatial relations. Linear probes on frozen SigLIP sit at 0.51 to 0.53 on them at the 256-token budget the baselines run, and more tokens do not close the gap.
  4. Feb 2026
    NVIDIA · arXiv:2503.14734 · archeologist
    Strongest in data engineering and as the first open humanoid VLA. The dual-system framing is more marketing than architecture, and N1.5 and N1.6 walked away from it within nine months.
  5. Feb 2026
    NVIDIA · arXiv:2511.00062 · reviewer
    A useful open release judged on its authors’ own benchmark, with a three-trial robotics experiment and no ablation of the recipe. Weak accept.
  6. Feb 2026
    Physical Intelligence · arXiv:2504.16054 · reviewer
    The first VLA shown working in homes it never saw. Accept, with one standard benchmark added and the PaliGemma backbone moved out of the appendix.
  7. Jan 2026
    FAST, Physical Intelligence · arXiv:2501.09747 · empiricist
    On Berkeley cable routing, FAST+ keeps a median 81% of a custom-fit tokenizer’s compression, right at my threshold, and the answer turns on how the custom tokenizer is tuned.