Does the universal FAST+ tokenizer transfer to cable routing?

A small experiment with Physical Intelligence’s FAST action tokenizer on real cable-routing data. The universal FAST+ tokenizer keeps a median 81% of a custom-fit tokenizer’s compression, right at my 80% threshold, and the answer turns on how the custom tokenizer is tuned.

Paper: FAST, arXiv:2501.09747 Tokenizer: physical-intelligence/fast Role: empiricist January 2026

Hypothesis

The FAST+ universal tokenizer achieves at least 80% of the compression ratio of a custom FAST tokenizer on cable manipulation data.

FAST compresses action chunks with a discrete cosine transform followed by byte-pair encoding, and FAST+ is the universal tokenizer trained across many robot datasets. The paper's pitch is that FAST+ works out of the box. Cable routing is a good stress test because it is contact-rich, low-rate and unlike most of the tabletop data FAST+ was fit on. The question matters for the Intrinsic AI for Industry challenge, where the task is cable assembly.

Setup

  • Data: lerobot/berkeley_cable_routing, 1,647 real episodes of a 7-DoF Franka arm routing a cable through clips at 10 Hz.
  • Tokenizers: naive per-dimension binning as in RT-2 and OpenVLA, the official FAST+ processor, and a custom FAST tokenizer fit on this dataset with the official .fit() on the same code path. Nothing is reimplemented locally.
  • Sweep: chunk length of 1 or 2 seconds, stride of 1, 2, 5 or 10 steps, BPE vocabulary of 1,024, 2,048 or 4,096, and DCT scale of 8, 10 or 12, for 72 configurations.
  • Environments: two, because OpenVLA's tokenizer pins draccus==0.8.0 and LeRobot pins 0.10.0. One environment loads the data, the other tokenizes.

Results

ChunkNaive tokensFAST+ compressionCustom FAST, medianCustom FAST, bestFAST+ ÷ custom, median
1 s (10 steps × 7 dims)702.60×3.38×4.65×0.77
2 s (20 steps × 7 dims)1402.78×3.32×4.72×0.84

Across all 72 configurations the FAST+ compression ratio stays in a narrow band, 2.58 to 2.79×, while the custom tokenizer ranges from 2.18 to 4.72×. The ratio of the two, which is what the hypothesis is about, runs from 0.56 to 1.19 with a median of 0.81. So the hypothesis holds at the median, by one point, and does not hold robustly. The deciding variable is the custom tokenizer's DCT scale. At a loose scale of 10 the custom tokenizer is worse than FAST+ on every vocabulary, so the ratio exceeds 1. At a tight scale of 8 with a 4,096 vocabulary it beats FAST+ by nearly 2×. Longer chunks help both tokenizers a little.

Both tokenizers compress far less than the 3 to 10× the paper reports on its datasets. My reading is that 10 Hz cable-routing actions carry less temporal redundancy than the 50 Hz data FAST was designed for, so there is less for the DCT to remove.

Two line charts. Left: next-token cross-entropy loss against sample length for naive, FAST+ and custom FAST tokens; naive stays near 2.9 while FAST+ and custom FAST sit between 4.5 and 5.5. Right: token prediction accuracy; naive stays near 0.46 while FAST tokens fall between 0.07 and 0.17.
A second metric, after the paper's Figure 3. Naive binned tokens are far easier to predict than FAST tokens at every chunk length.

The second metric reproduces the paper's motivation rather than its headline. Naive tokens are almost trivially predictable from their predecessors, with cross-entropy near 2.9 and accuracy near 0.46 regardless of chunk length, because consecutive per-dimension bins are nearly identical at any reasonable control rate. FAST tokens are much harder to predict, with loss between 4.5 and 5.5 and accuracy below 0.17. That is the point of the DCT: it removes the redundancy that would otherwise let a policy minimize its loss without learning anything about the task.

What I take from it

  • FAST+ is a safe default on out-of-distribution data: a stable 2.6 to 2.8× with no fitting step.
  • A custom tokenizer is worth it only if you tune it. Default-ish settings can be worse than FAST+, and the best settings nearly double its compression.
  • Compression is a proxy. The experiment I would run next is to train a small policy on each tokenization of the same demonstrations and measure downstream success, then repeat at 50 Hz where FAST's advantage should be larger.

Method. Scripts, sweep table and figure are mine, from January 2026. All tokenization uses the official processor from Hugging Face.