Use the chapter buttons to jump directly to each part of the overview video.
A vision-language-action policy may understand the task correctly while the robot cannot execute its intended motion. If a joint locks, the same action command produces a different physical response. CAPABLE estimates that changing capability from the robot's own command–response history and adjusts a frozen VLA policy with a bounded residual action.
No fault label, affected-joint identifier, or fault-specific demonstration is needed at deployment. The nominal VLA continues to provide task guidance while the residual controller adapts the arm's motion.
A shared temporal encoder processes the previous 16 command–response observations independently for each of the Franka Panda's seven joints. The live Jacobian grounds each joint representation in its current contribution to end-effector motion. Cross-joint attention combines these representations into a capability latent.
Self-supervised prediction of realized joint and end-effector motion shapes the latent. A FiLM-conditioned SAC residual actor then corrects the six arm action dimensions proposed by the frozen OpenVLA-OFT backbone. The gripper command passes through unchanged.
The principal experiment trains with persistent locks on j₀, j₄, j₅, and j₆; joint j₂ is withheld from every training stage and appears only in evaluation.
Across 28 LIBERO tasks, CAPABLE raises success on the globally unseen j₂ lock from 24.8% for the frozen VLA to 59.3%. A parameter-matched global-history SAC residual baseline reaches 41.9%, a 17.4-point CAPABLE advantage. Healthy success remains 90.8% versus 91.4% for the frozen VLA.
| Method | Healthy | Seen locks | Unseen j₂ lock |
|---|---|---|---|
| Base VLA | 91.4 | 41.6 | 24.8 |
| Classical redundancy resolution | 81.6 | 61.3 | 32.9 |
| DEFT (privileged) | 89.7 ± 1.4 | 70.1 ± 2.3 | 62.8 ± 2.9 |
| Global-history SAC | 88.1 ± 1.3 | 73.8 ± 2.1 | 41.9 ± 3.2 |
| CAPABLE | 90.8 ± 1.0 | 86.7 ± 1.8 | 59.3 ± 3.0 |
Task-balanced success (%). ± is standard deviation across three training seeds for learned methods. DEFT receives privileged actuator availability and simulator-derived goals. Global-history SAC is the matched representation comparison.
On six independent leave-one-actuator-out splits over a separate eight-task subset, CAPABLE leads the matched baseline on every held-out joint; mean success is 67.0% versus 42.4%. Lock-trained checkpoints also improve on unseen-j₂ damping (+12.2 points), friction (+12.4), and late-onset locks (+10.9) over the matched baseline. Under partial effectiveness and range restriction, the matched baseline leads by 1.8 and 6.5 points, respectively.
CAPABLE and Global-history SAC are transferred from a digital twin to a physical Franka Panda without hardware RL. A joint lock is software-enforced at the joint-reference layer. Each task–joint condition has ten trials per method.
| Task | Joint | Global-history SAC | CAPABLE |
|---|---|---|---|
| Open drawer | j₀ (seen) | 5/10 | 9/10 |
| Put object in drawer | j₂ (unseen) | 4/10 | 7/10 |
| Object rearrangement | j₆ (seen) | 4/10 | 8/10 |
| Overall | — | 13/30 (43.3%) | 24/30 (80.0%) |