Projects

Code and write-ups. Each links to its page or repository.

  1. Benchmark
    A long-horizon benchmark built on NVIDIA Isaac Lab. In one episode a Unitree G1 positions a step ladder under a fixture, climbs it, exchanges a spent bulb for a fresh one and disposes of the spent one. Twelve subtask environments scored on difficulty-weighted gates, standard and privileged observation modes, and reference baselines spanning RSL-RL PPO, zero-shot GR00T N1.7 and whole-body controllers.
  2. Analysis
    What the 100-task BEHAVIOR Challenge 2026 benchmark is made of, read from the evaluator’s code and all 20,000 recorded demonstrations.
  3. Reproduction
    Priced the report’s five claims against available compute, characterised the released rejection-sampling dataset without a GPU, and evaluated the released π0.5-based checkpoint on two machines. Fixes to the evaluation stack are in the fork.
  4. Experiment
    Linear probes on patch tokens with exact labels replayed from the simulator, a weights-pinned token-budget sweep with SigLIP 2 NaFlex, and an encoder swap against SigLIP 1 and DINOv2. Preliminary results sit near chance on spatial relations.
  5. Workshop
    A pipeline that turns an ordinary video into motion on a Unitree G1 humanoid: PromptHMR estimates 3D pose, GMR retargets the joints, and GR00T whole-body control drives simulation or hardware. The codebase for the ʻIolani workshop on applied ML and humanoid robots.
  6. Experiments
    Three experiments on adding 6-axis force-torque sensing to SmolVLA, a 0.45B vision-language-action model, for contact-rich manipulation tasks. SmolVLA takes vision and joint positions as input but has no sense of touch; the experiments test whether force augmentation works at that scale.
  7. Experiment
    Physical Intelligence’s FAST+ action tokenizer against a custom-fit FAST tokenizer and naive binning on 1,647 real Berkeley cable-routing episodes, across 72 configurations. FAST+ keeps a median 81% of the custom tokenizer’s compression.