Related Work: Force-Torque in Vision-Language-Action Models

A survey of how the field is approaching force/torque integration into VLAs, organized by approach cluster.

Landscape

Paper Date Approach Base Model Key Result
ForceVLA May 2025 Force-aware MoE fusion pi0 (3.3B) +23.2% on contact-rich tasks, 80% plug insertion
FTACT Sep 2025 F/T-aware ACT ACT (80M) 60% -> 80% on bottle reorientation
VLA-Touch Jul 2025 Dual-level tactile adapter Any VLA No base VLA fine-tuning needed
CRAFT 2025 Cross-modal attention F/T fusion Custom Contact-rich task generalization
TaF-VLA 2025 Tactile-force VLA – Combined tactile + F/T sensing
ForceMimic 2025 Force-aware imitation learning IL baseline Force profile matching in demos
Direction Matters 2025 Directional F/T encoding – Orientation-aware force representation
FoAR 2025 Force-aware residual policy RL + residual F/T-conditioned corrections
FILIC 2025 Force-informed learned impedance – Compliance from F/T feedback
CompliantVLA-adaptor Jan 2026 VLM-guided impedance Any VLA 9.86% -> 17.29% avg (safety-focused)
FD-VLA Feb 2026 Force distillation from vision pi0 Predicts F/T without a sensor
TA-VLA 2025 Torque integration design – Torque in decoder, not encoder
HumanoidVLM Jan 2026 Impedance control – Humanoid contact manipulation

Approach Clusters

1. F/T as Input Modality

Papers: ForceVLA, FTACT, CRAFT, TaF-VLA

Directly feed wrench data into the policy. ForceVLA’s MoE (Mixture of Experts) approach is the most mature, achieving +23.2% on contact-rich tasks. The challenge is where to inject the signal – FTACT feeds it into the encoder, while TA-VLA argues torque should go to the decoder.

Our Experiment 1 tests this approach at SmolVLA scale.

2. Force Distillation

Papers: FD-VLA

Learn force representations from vision alone – deploy without sensors. FD-VLA trains an auxiliary force prediction head during training, then discards it at inference. The VLM features carry implicit force understanding.

Our Experiment 2 tests whether this compresses from pi0 (3.3B) to SmolVLA (0.45B).

3. Compliance / Impedance Control

Papers: CompliantVLA-adaptor, FILIC, HumanoidVLM

F/T drives variable stiffness, not just actions. The VLA outputs impedance parameters (stiffness, damping) rather than position targets. This is inherently safer for contact-rich tasks.

4. Residual Correction

Papers: FoAR

F/T drives a fast corrective layer on top of a base policy. The base policy handles high-level planning; the residual handles force-reactive fine adjustments.

Our Experiment 4 implements this dual-rate approach with SmolVLA as the base policy.

5. Representation Design

Papers: TA-VLA, Direction Matters

Where and how force enters the architecture matters:

  • TA-VLA: Torque should go to the decoder (action expert), not the encoder (VLM). SmolVLA’s own ablation says states work better as VLM prefix.
  • Direction Matters: Directional encoding of F/T (preserving orientation) outperforms raw scalar values. Encode wrench in tool frame, not base frame.

6. Tactile Fusion

Papers: VLA-Touch, TaF-VLA, ForceMimic

Broader touch modalities beyond wrist F/T – skin sensors, fingertip arrays. VLA-Touch’s dual-level adapter is notable for requiring no base VLA fine-tuning.

The Gap We Address

All force-VLA work targets pi0 (3.3B) or ACT (80M). Nobody has tested force augmentation on an efficient, community-pretrained VLA like SmolVLA (0.45B). The efficiency question – does force-torque integration work at 10x smaller scale? – was unanswered until these experiments.


SmolVLA Force-Torque Experiments — Pavel Bushuyeu, 2026

This site uses Just the Docs, a documentation theme for Jekyll.