SmolVLA Force-Torque Experiments
Three experiments on adding 6-axis force-torque (F/T) sensing to SmolVLA, a 0.45B Vision-Language-Action model from HuggingFace.
Motivation
SmolVLA takes vision + joint positions as input. For contact-rich tasks like cable insertion, the robot needs to feel – detecting whether a connector is seated, whether it’s pushing too hard, or slipping sideways. A force-torque sensor on the robot’s wrist provides this: 6 numbers (3 forces + 3 torques) updated hundreds of times per second.
All prior work on force-augmented VLAs targets larger models: ForceVLA and FD-VLA use pi0 (3.3B), FTACT uses ACT (80M). Nobody has tested force augmentation on an efficient VLA at 0.45B scale. These experiments fill that gap.
Dataset & Hardware
- Dataset: RH20T – 30 episodes of a UR5 robot arm, 4,933 frames, with synchronized camera images, joint positions, and 6-axis F/T readings
- Hardware: NVIDIA GH200 (96GB VRAM), Lambda Cloud
- Base model: SmolVLA 0.45B (
lerobot/smolvla_base) – 450M total params, ~100M trainable
Results
| Experiment | Approach | Added Params | Training Time | Key Finding |
|---|---|---|---|---|
| Exp 1: Force Token | Append F/T to state vector | ~0 | 5h/variant (25h total) | Null result – all variants converge identically |
| Exp 2: Force Distill | Auxiliary force prediction head | +75K | 3-8h/config | Tests “force imagination” from vision at small scale |
| Exp 4: Dual-Rate | Separate 500Hz force corrector MLP | 5,382 | 2 seconds | Near-zero val loss; most deployment-ready |
Limitations
These are offline experiments with significant caveats:
- No real images – RH20T on HuggingFace lacks image columns; blank images were injected, so the VLM backbone learned nothing from vision
- Tiny F/T values – forces peak at ~1.3N (light contact); cable insertion tasks involve 10-50N
- Small dataset – 4,933 frames / 30 episodes
- No real-robot evaluation – all metrics are offline (action prediction loss)
Next Steps
Apply these approaches to the Intrinsic AI Challenge (UR5e + Axia80 F/T + 3 cameras) once the simulation toolkit is available. The dual-rate hybrid (Exp 4) is the most deployment-ready approach.