Experiment 2: Force Distillation
Hypothesis: SmolVLA can learn to predict force-torque from visual observations alone (force distillation), and this internal force representation improves action quality even without a physical F/T sensor at inference time.
Inspired by FD-VLA (Feb 2026), which demonstrated this on pi0 (3.3B). We test whether it compresses to 0.45B.
The Idea
Instead of feeding F/T to the model, teach it to predict forces from what it sees. A small “force prediction head” branches off the VLM and tries to guess the F/T readings. The training loss becomes:
total_loss = action_loss + lambda * force_prediction_loss
At inference time, the force head is discarded – but the VLM has learned internal representations that encode force information. It has learned to “imagine” forces from visual cues.
Architecture
+----------------+
| Force Head | <- MLP (576 -> 128 -> 128 -> 6)
| (aux loss) | supervised by real F/T from RH20T
+-------+--------+
| taps VLM hidden states (mean pooled)
+---------------------------+-----------------------------+
| +------+-----+ +---------+ |
| | |----->| | |
| | | kv | | |
| | |----->| Action | |
| | VLM |cache | Expert | |
| | |----->| | |
| | | | | |
| +^--^----^---+ +----^----+ |
| | | | | |
| | | state noise |
| | language |
| images |
+--------------------------------------------------------+
Force head: MLP (576 -> 128 -> 128 -> 6) – takes mean-pooled VLM features, predicts 6-axis wrench. Only ~75K parameters added to the 450M base model.
Configuration Sweep
A 2x3 sweep over lambda (force loss weight) and VLM backbone freezing:
| Config | Lambda | VLM Backbone | Steps | Trainable Params |
|---|---|---|---|---|
lambda_0.1_frozen | 0.1 | Frozen | 30,000 | 100M |
lambda_0.1_unfrozen | 0.1 | Last 4 layers unfrozen | 10,000 | 139M |
lambda_0.2_frozen | 0.2 | Frozen | 10,000 | 100M |
lambda_0.2_unfrozen | 0.2 | Last 4 layers unfrozen | 10,000 | 139M |
lambda_0.5_frozen | 0.5 | Frozen | 10,000 | 100M |
lambda_0.5_unfrozen | 0.5 | Last 4 layers unfrozen | 10,000 | 139M |
Analysis
The key question: does the auxiliary force objective improve action prediction quality?
- Frozen VLM: The force head must extract force information from existing VLM features. If it succeeds, force-relevant information is already present in the representations – distillation just makes it explicit.
- Unfrozen VLM: The backbone can reorganize its representations to better serve both action prediction and force prediction. This tests whether “force-aware” features need to permeate the VLM.
The frozen vs unfrozen comparison tells us whether force-aware representations need to permeate the VLM backbone or can be extracted from existing features.
Training Details
- Base model: SmolVLA 0.45B (
lerobot/smolvla_base) - Force pool mode: Mean pooling over VLM sequence features
- Optimizer: AdamW, LR 1e-4, linear warmup (1K steps) + cosine decay
- Batch size: 64
- Gradient clipping: 1.0
- Hardware: NVIDIA GH200 96GB, Lambda Cloud
- WandB:
smolvla-force-distill
Code
See experiments/exp2_force_distill/ for the full implementation.
Key files:
modeling.py– patched SmolVLA with force prediction headtrain.py– training loop with auxiliary force lossdata_pipeline.py– RH20T to LeRobot format converterevaluate.py– three-way evaluation: baseline vs real-FT vs distilledconfigs/– YAML configs for lambda sweep