Experiment 2: Force Distillation

Hypothesis: SmolVLA can learn to predict force-torque from visual observations alone (force distillation), and this internal force representation improves action quality even without a physical F/T sensor at inference time.

Inspired by FD-VLA (Feb 2026), which demonstrated this on pi0 (3.3B). We test whether it compresses to 0.45B.

The Idea

Instead of feeding F/T to the model, teach it to predict forces from what it sees. A small “force prediction head” branches off the VLM and tries to guess the F/T readings. The training loss becomes:

total_loss = action_loss + lambda * force_prediction_loss

At inference time, the force head is discarded – but the VLM has learned internal representations that encode force information. It has learned to “imagine” forces from visual cues.

Architecture

                    +----------------+
                    |  Force Head    |  <- MLP (576 -> 128 -> 128 -> 6)
                    |  (aux loss)    |     supervised by real F/T from RH20T
                    +-------+--------+
                            | taps VLM hidden states (mean pooled)
+---------------------------+-----------------------------+
|                    +------+-----+      +---------+      |
|                    |            |----->|         |      |
|                    |            | kv   |         |      |
|                    |            |----->| Action  |      |
|                    |   VLM      |cache | Expert  |      |
|                    |            |----->|         |      |
|                    |            |      |         |      |
|                    +^--^----^---+      +----^----+      |
|                     |  |    |              |            |
|                     |  |    state        noise          |
|                     |  language                         |
|                     images                              |
+--------------------------------------------------------+

Force head: MLP (576 -> 128 -> 128 -> 6) – takes mean-pooled VLM features, predicts 6-axis wrench. Only ~75K parameters added to the 450M base model.

Configuration Sweep

A 2x3 sweep over lambda (force loss weight) and VLM backbone freezing:

Config Lambda VLM Backbone Steps Trainable Params
lambda_0.1_frozen 0.1 Frozen 30,000 100M
lambda_0.1_unfrozen 0.1 Last 4 layers unfrozen 10,000 139M
lambda_0.2_frozen 0.2 Frozen 10,000 100M
lambda_0.2_unfrozen 0.2 Last 4 layers unfrozen 10,000 139M
lambda_0.5_frozen 0.5 Frozen 10,000 100M
lambda_0.5_unfrozen 0.5 Last 4 layers unfrozen 10,000 139M

Analysis

The key question: does the auxiliary force objective improve action prediction quality?

  • Frozen VLM: The force head must extract force information from existing VLM features. If it succeeds, force-relevant information is already present in the representations – distillation just makes it explicit.
  • Unfrozen VLM: The backbone can reorganize its representations to better serve both action prediction and force prediction. This tests whether “force-aware” features need to permeate the VLM.

The frozen vs unfrozen comparison tells us whether force-aware representations need to permeate the VLM backbone or can be extracted from existing features.

Training Details

  • Base model: SmolVLA 0.45B (lerobot/smolvla_base)
  • Force pool mode: Mean pooling over VLM sequence features
  • Optimizer: AdamW, LR 1e-4, linear warmup (1K steps) + cosine decay
  • Batch size: 64
  • Gradient clipping: 1.0
  • Hardware: NVIDIA GH200 96GB, Lambda Cloud
  • WandB: smolvla-force-distill

Code

See experiments/exp2_force_distill/ for the full implementation.

Key files:

  • modeling.py – patched SmolVLA with force prediction head
  • train.py – training loop with auxiliary force loss
  • data_pipeline.py – RH20T to LeRobot format converter
  • evaluate.py – three-way evaluation: baseline vs real-FT vs distilled
  • configs/ – YAML configs for lambda sweep

SmolVLA Force-Torque Experiments — Pavel Bushuyeu, 2026

This site uses Just the Docs, a documentation theme for Jekyll.