CAPABLE

Capability-Aware Policy Adaptation via Behavioral Latent Encoding
University of BremenUniversity of North TexasToyota Motor North America

Overview

Use the chapter buttons to jump directly to each part of the overview video.

Why capability-aware policy adaptation?

A vision-language-action policy may understand the task correctly while the robot cannot execute its intended motion. If a joint locks, the same action command produces a different physical response. CAPABLE estimates that changing capability from the robot's own command–response history and adjusts a frozen VLA policy with a bounded residual action.

No fault label, affected-joint identifier, or fault-specific demonstration is needed at deployment. The nominal VLA continues to provide task guidance while the residual controller adapts the arm's motion.

CAPABLE paper overview showing command-response history and Jacobian, capability inference, a residual RL actor, a frozen OpenVLA-OFT policy, and a Franka robot with an impaired joint.
CAPABLE paper overview. Joint behavior and live kinematics condition a residual action on top of a frozen VLA policy.

How does CAPABLE work?

A shared temporal encoder processes the previous 16 command–response observations independently for each of the Franka Panda's seven joints. The live Jacobian grounds each joint representation in its current contribution to end-effector motion. Cross-joint attention combines these representations into a capability latent.

Self-supervised prediction of realized joint and end-effector motion shapes the latent. A FiLM-conditioned SAC residual actor then corrects the six arm action dimensions proposed by the frozen OpenVLA-OFT backbone. The gripper command passes through unchanged.

executed arm action = frozen VLA action + bounded capability-aware residual

The principal experiment trains with persistent locks on j₀, j₄, j₅, and j₆; joint j₂ is withheld from every training stage and appears only in evaluation.

Held-out joint recovery on LIBERO

Across 28 LIBERO tasks, CAPABLE raises success on the globally unseen j₂ lock from 24.8% for the frozen VLA to 59.3%. A parameter-matched global-history SAC residual baseline reaches 41.9%, a 17.4-point CAPABLE advantage. Healthy success remains 90.8% versus 91.4% for the frozen VLA.

MethodHealthySeen locksUnseen j₂ lock
Base VLA91.441.624.8
Classical redundancy resolution81.661.332.9
DEFT (privileged)89.7 ± 1.470.1 ± 2.362.8 ± 2.9
Global-history SAC88.1 ± 1.373.8 ± 2.141.9 ± 3.2
CAPABLE90.8 ± 1.086.7 ± 1.859.3 ± 3.0

Task-balanced success (%). ± is standard deviation across three training seeds for learned methods. DEFT receives privileged actuator availability and simulator-derived goals. Global-history SAC is the matched representation comparison.

Transfer beyond one held-out joint

On six independent leave-one-actuator-out splits over a separate eight-task subset, CAPABLE leads the matched baseline on every held-out joint; mean success is 67.0% versus 42.4%. Lock-trained checkpoints also improve on unseen-j₂ damping (+12.2 points), friction (+12.4), and late-onset locks (+10.9) over the matched baseline. Under partial effectiveness and range restriction, the matched baseline leads by 1.8 and 6.5 points, respectively.

Real-robot evaluation

CAPABLE and Global-history SAC are transferred from a digital twin to a physical Franka Panda without hardware RL. A joint lock is software-enforced at the joint-reference layer. Each task–joint condition has ten trials per method.

TaskJointGlobal-history SACCAPABLE
Open drawerj₀ (seen)5/109/10
Put object in drawerj₂ (unseen)4/107/10
Object rearrangementj₆ (seen)4/108/10
Overall—13/30 (43.3%)24/30 (80.0%)