Related Work: Force-Torque in Vision-Language-Action Models
A survey of how the field is approaching force/torque integration into VLAs, organized by approach cluster.
Landscape
| Paper | Date | Approach | Base Model | Key Result |
|---|---|---|---|---|
| ForceVLA | May 2025 | Force-aware MoE fusion | pi0 (3.3B) | +23.2% on contact-rich tasks, 80% plug insertion |
| FTACT | Sep 2025 | F/T-aware ACT | ACT (80M) | 60% -> 80% on bottle reorientation |
| VLA-Touch | Jul 2025 | Dual-level tactile adapter | Any VLA | No base VLA fine-tuning needed |
| CRAFT | 2025 | Cross-modal attention F/T fusion | Custom | Contact-rich task generalization |
| TaF-VLA | 2025 | Tactile-force VLA | – | Combined tactile + F/T sensing |
| ForceMimic | 2025 | Force-aware imitation learning | IL baseline | Force profile matching in demos |
| Direction Matters | 2025 | Directional F/T encoding | – | Orientation-aware force representation |
| FoAR | 2025 | Force-aware residual policy | RL + residual | F/T-conditioned corrections |
| FILIC | 2025 | Force-informed learned impedance | – | Compliance from F/T feedback |
| CompliantVLA-adaptor | Jan 2026 | VLM-guided impedance | Any VLA | 9.86% -> 17.29% avg (safety-focused) |
| FD-VLA | Feb 2026 | Force distillation from vision | pi0 | Predicts F/T without a sensor |
| TA-VLA | 2025 | Torque integration design | – | Torque in decoder, not encoder |
| HumanoidVLM | Jan 2026 | Impedance control | – | Humanoid contact manipulation |
Approach Clusters
1. F/T as Input Modality
Papers: ForceVLA, FTACT, CRAFT, TaF-VLA
Directly feed wrench data into the policy. ForceVLA’s MoE (Mixture of Experts) approach is the most mature, achieving +23.2% on contact-rich tasks. The challenge is where to inject the signal – FTACT feeds it into the encoder, while TA-VLA argues torque should go to the decoder.
Our Experiment 1 tests this approach at SmolVLA scale.
2. Force Distillation
Papers: FD-VLA
Learn force representations from vision alone – deploy without sensors. FD-VLA trains an auxiliary force prediction head during training, then discards it at inference. The VLM features carry implicit force understanding.
Our Experiment 2 tests whether this compresses from pi0 (3.3B) to SmolVLA (0.45B).
3. Compliance / Impedance Control
Papers: CompliantVLA-adaptor, FILIC, HumanoidVLM
F/T drives variable stiffness, not just actions. The VLA outputs impedance parameters (stiffness, damping) rather than position targets. This is inherently safer for contact-rich tasks.
4. Residual Correction
Papers: FoAR
F/T drives a fast corrective layer on top of a base policy. The base policy handles high-level planning; the residual handles force-reactive fine adjustments.
Our Experiment 4 implements this dual-rate approach with SmolVLA as the base policy.
5. Representation Design
Papers: TA-VLA, Direction Matters
Where and how force enters the architecture matters:
- TA-VLA: Torque should go to the decoder (action expert), not the encoder (VLM). SmolVLA’s own ablation says states work better as VLM prefix.
- Direction Matters: Directional encoding of F/T (preserving orientation) outperforms raw scalar values. Encode wrench in tool frame, not base frame.
6. Tactile Fusion
Papers: VLA-Touch, TaF-VLA, ForceMimic
Broader touch modalities beyond wrist F/T – skin sensors, fingertip arrays. VLA-Touch’s dual-level adapter is notable for requiring no base VLA fine-tuning.
The Gap We Address
All force-VLA work targets pi0 (3.3B) or ACT (80M). Nobody has tested force augmentation on an efficient, community-pretrained VLA like SmolVLA (0.45B). The efficiency question – does force-torque integration work at 10x smaller scale? – was unanswered until these experiments.