Experiment 4: Dual-Rate Hybrid Controller

Hypothesis: A two-tier controller where SmolVLA generates position targets at ~30Hz and a lightweight force-conditioned corrector operates at 500Hz outperforms either approach alone for insertion tasks.

The Idea

Instead of modifying SmolVLA at all, train a tiny separate network as a “force correction layer”:

  1. Tier 1 – SmolVLA (slow, smart): Looks at cameras and says “move the arm 5cm to the left” (runs at ~30Hz)
  2. Tier 2 – Force Corrector (fast, reactive): Reads the F/T sensor and says “but you’re pushing too hard, ease off by 0.5mm” (runs at 500Hz)

The corrector takes 12 inputs (6-axis F/T + 6-axis position error) and outputs 6 position corrections. Blended output:

final_target = smolvla_target + alpha * corrector_output

Architecture

                       ~30Hz                          500Hz
  +------------+    +----------+    +--------------+    +---------+
  | 3 Cameras  |--->| SmolVLA  |--->| Action Chunk |--->|         |
  | Language   |    | (0.45B)  |    | (50 steps)   |    |  Blend  |---> UR5e
  | Joint State|    +----------+    +------+-------+    |         |    RTDE
  +------------+                           |            | target  |
                                     current target     |   +     |
                                           |            | alpha * |
                                           v            |  corr   |
  +------------+    +----------+    +--------------+    |         |
  | Axia80 F/T |--->| Tier 2   |--->|  Position    |--->|         |
  | (6-axis)   |    | MLP      |    |  Correction  |    +---------+
  | Pos Error  |--->| (~5K p.) |    |  (6 DoF)     |
  +------------+    +----------+    +--------------+

Tier 2 MLP:

Input (12) -> Linear(64) -> ReLU -> Linear(64) -> ReLU -> Linear(6) -> Output

Total parameters: 5,382 (compared to SmolVLA’s 450 million).

Results

Metric Value
Training samples 1,671
Validation samples 185
Best validation loss 1.05e-5
Training time 2.0 seconds
Epochs 50
Contact threshold 0.05N

Analysis

The corrector trained in 2 seconds and achieved extremely low validation loss. This is expected – the mapping from (force, position_error) to corrections is relatively simple and well-structured.

Why this is the most deployment-ready approach:

  • No VLA modification required – SmolVLA runs unmodified; the corrector is a separate module
  • Trivial to train – 5K parameters, 2 seconds, no GPU needed
  • Real-time capable – the MLP runs at 500Hz on a CPU, matching UR5e’s RTDE interface
  • Safety built in – the corrector has force limits, correction clipping, and emergency stop logic
  • Composable – works with any slow policy, not just SmolVLA

The real test would be deploying this alongside SmolVLA on a real robot, where the corrector provides sub-millisecond force-reactive adjustments while SmolVLA handles high-level planning.

This approach is inspired by FoAR (Force-aware Residual policy, 2025) and directly applicable to the Intrinsic AI Challenge’s UR5e, which supports 500Hz real-time control.

Configuration Variants

Config Alpha Max Correction Contact Threshold
default.yaml 0.3 5mm 0.05N
aggressive_correction.yaml 0.6 10mm lower
conservative_correction.yaml 0.15 2mm tighter force limits

Training Details

  • Optimizer: Adam, LR 1e-3, cosine schedule
  • F/T normalization: Both inputs and correction targets normalized
  • Contact threshold: 0.05N (values in dataset max at ~1.3N)
  • WandB: smolvla-dual-rate

Code

See experiments/exp4_dual_rate/ for the full implementation.

Key files:

  • tier1_smolvla.py – async SmolVLA inference wrapper
  • tier2_corrector.py – force residual MLP (12 -> 64 -> 64 -> 6)
  • data_extraction.py – extract (F/T, pos_error) -> correction pairs from RH20T
  • train_tier2.py – supervised MLP training
  • controller.py – integration loop with safety (force limits, clipping, e-stop)
  • evaluate.py – compare SmolVLA-alone vs corrector-alone vs hybrid

SmolVLA Force-Torque Experiments — Pavel Bushuyeu, 2026

This site uses Just the Docs, a documentation theme for Jekyll.