Experiment 1: Force Token Injection

Hypothesis: Concatenating 6-axis F/T readings to SmolVLA’s sensorimotor state vector improves action prediction on contact-rich tasks, with minimal parameter overhead.

The Idea

SmolVLA reads joint positions (6 numbers) through a “state projector” – a small linear layer (nn.Linear(32, 960)). We append the F/T readings to the joint positions before feeding them into this layer. It’s the simplest possible integration: give the model a sticky note that says “here’s what my wrist feels” alongside the usual joint angles.

Variants

Five variants, each trained for 20,000 steps (~260 epochs):

Variant State Vector Contents Extra Dims
baseline Joint positions only (6 dims, padded to 32) 0
force_only Joints + force XYZ (3 dims) 3
torque_only Joints + torque XYZ (3 dims) 3
wrench_packed Joints + full wrench (6 dims), packed into existing 32-dim slot 6
wrench_extended Joints + full wrench, state projector expanded from 32 to 38 dims 6

Results

Variant Loss @ 1K Loss @ 5K Loss @ 10K Loss @ 20K (final)
baseline 0.004885 0.002105 0.001026 0.000320
force_only 0.004788 0.002061 0.001032 0.000319
torque_only 0.005085 0.002082 0.001032 0.000318
wrench_packed 0.004842 0.002113 0.001028 0.000317
wrench_extended 0.005615 0.001994 0.001015 0.000316

Analysis

This is a null result. All five variants converged to nearly identical final losses (~0.000316-0.000320). The wrench_extended variant achieved the lowest final loss, but the differences are within noise.

Why no effect?

  • No real images: The RH20T HuggingFace dataset lacks image columns. We injected blank dummy images, so the VLM backbone learned nothing from vision. The model is essentially just a state-to-action MLP.
  • Tiny F/T values: Forces in this dataset peak at ~1.3N (light contact). For heavy-contact tasks like cable insertion (10-50N), F/T information would be far more discriminative.
  • Small dataset: 4,903 samples / 30 episodes. Not enough diversity for the model to learn distinct F/T-conditioned behaviors.

The null result doesn’t mean F/T is useless for small VLAs – it means this particular dataset doesn’t have enough signal. On contact-rich tasks with real images and higher forces, the result could be very different.

Training Details

  • Model: SmolVLA 0.45B (450M total, ~100M trainable)
  • State projector: nn.Linear(32, 960) (or nn.Linear(38, 960) for extended)
  • Optimizer: AdamW, cosine LR schedule, peak LR 1e-4
  • Hardware: NVIDIA GH200 96GB, Lambda Cloud
  • WandB: smolvla-force-token

Code

See experiments/exp1_force_token/ for the full implementation.

Key files:

  • scripts/train_ft.py – training loop
  • patches/patch_smolvla.py – patches SmolVLA to accept extended state
  • patches/rh20t_ft_dataset.py – RH20T dataset with F/T columns
  • configs/ – YAML configs for each variant

SmolVLA Force-Torque Experiments — Pavel Bushuyeu, 2026

This site uses Just the Docs, a documentation theme for Jekyll.