Experiment 1: Force Token Injection
Hypothesis: Concatenating 6-axis F/T readings to SmolVLA’s sensorimotor state vector improves action prediction on contact-rich tasks, with minimal parameter overhead.
The Idea
SmolVLA reads joint positions (6 numbers) through a “state projector” – a small linear layer (nn.Linear(32, 960)). We append the F/T readings to the joint positions before feeding them into this layer. It’s the simplest possible integration: give the model a sticky note that says “here’s what my wrist feels” alongside the usual joint angles.
Variants
Five variants, each trained for 20,000 steps (~260 epochs):
| Variant | State Vector Contents | Extra Dims |
|---|---|---|
| baseline | Joint positions only (6 dims, padded to 32) | 0 |
| force_only | Joints + force XYZ (3 dims) | 3 |
| torque_only | Joints + torque XYZ (3 dims) | 3 |
| wrench_packed | Joints + full wrench (6 dims), packed into existing 32-dim slot | 6 |
| wrench_extended | Joints + full wrench, state projector expanded from 32 to 38 dims | 6 |
Results
| Variant | Loss @ 1K | Loss @ 5K | Loss @ 10K | Loss @ 20K (final) |
|---|---|---|---|---|
| baseline | 0.004885 | 0.002105 | 0.001026 | 0.000320 |
| force_only | 0.004788 | 0.002061 | 0.001032 | 0.000319 |
| torque_only | 0.005085 | 0.002082 | 0.001032 | 0.000318 |
| wrench_packed | 0.004842 | 0.002113 | 0.001028 | 0.000317 |
| wrench_extended | 0.005615 | 0.001994 | 0.001015 | 0.000316 |
Analysis
This is a null result. All five variants converged to nearly identical final losses (~0.000316-0.000320). The wrench_extended variant achieved the lowest final loss, but the differences are within noise.
Why no effect?
- No real images: The RH20T HuggingFace dataset lacks image columns. We injected blank dummy images, so the VLM backbone learned nothing from vision. The model is essentially just a state-to-action MLP.
- Tiny F/T values: Forces in this dataset peak at ~1.3N (light contact). For heavy-contact tasks like cable insertion (10-50N), F/T information would be far more discriminative.
- Small dataset: 4,903 samples / 30 episodes. Not enough diversity for the model to learn distinct F/T-conditioned behaviors.
The null result doesn’t mean F/T is useless for small VLAs – it means this particular dataset doesn’t have enough signal. On contact-rich tasks with real images and higher forces, the result could be very different.
Training Details
- Model: SmolVLA 0.45B (450M total, ~100M trainable)
- State projector:
nn.Linear(32, 960)(ornn.Linear(38, 960)for extended) - Optimizer: AdamW, cosine LR schedule, peak LR 1e-4
- Hardware: NVIDIA GH200 96GB, Lambda Cloud
- WandB:
smolvla-force-token
Code
See experiments/exp1_force_token/ for the full implementation.
Key files:
scripts/train_ft.py– training looppatches/patch_smolvla.py– patches SmolVLA to accept extended statepatches/rh20t_ft_dataset.py– RH20T dataset with F/T columnsconfigs/– YAML configs for each variant