Multi-Link Safety Filtering for VLA Policies Around Moving Hazards

Yatharth Agarwal Vijay Raghunathan

School of Electrical and Computer Engineering, Purdue University

No filter The arm knocks the bottle over.
With our filter The cube reaches the cup and the bottle stays upright.

A robot policy that follows language instructions can finish its task and still knock over things it was never asked to touch. We add a safety filter between the policy and the robot that keeps the whole arm away from a hazard, even while someone moves it, without retraining the policy.

The idea in three parts

1 Guard the whole arm

Many safety filters for these policies only guard the gripper. Ours also guards the wrist and forearm, which can hit things before the gripper does as the arm reaches across the table.

2 Follow moving objects

The object to avoid is found once, at the start. After that a lightweight tracker follows it as it moves, instead of searching for it again in every frame.

3 Run on one laptop

The robot policy, the object finder and the safety filter all run on a single laptop processor. The filter only nudges the policy's commands when the arm gets too close.

Moving hazards in simulation

We built a simulated test in which the object to avoid moves while the robot works.

Six hazard conditions: stationary, a 25 mm orbit, a 300 mm shuttle, and one-way escapes of 50, 150 and 300 mm Stationary Orbit 25 mm Shuttle 300 mm Escape 50 mm Escape 150 mm Escape 300 mm
Six hazard conditions: one stationary, five moving at 50 mm/s.
Simulator
MuJoCo · Franka Panda, LIBERO suites
Policy
π0.5, public LIBERO checkpoint
Perception
Rendered RGB and registered depth
Collision
Hazard displaced over 1 mm at any step

Averaged over all six:

Episodes that hit the object
65.62%27.27%
no filter → with our filter (lower is better)
Task done, object untouched
29.35%50.43%
no filter → with our filter (higher is better)

Each clip runs the same scene and start twice: left, no filter; right, with our filter.

Stationary. Without the filter, the forearm, not the gripper, strikes the moka pot.
Orbit 25 mm. The wine bottle circles in place.
Shuttle 300 mm. The book never stops moving.
Escape 50 mm. The milk carton slides a little, then stops.
Escape 150 mm. The carton slides farther before it stops.
Escape 300 mm. The carton travels the farthest, so tracking it matters most.

These are hand-picked examples in which the policy alone finishes the task but hits the object. The numbers above come from the full test in the paper.

On a real robot arm

The same filter runs on a low-cost SO-101 arm while a person carries a bottle into its path. The robot's task is to place a sugar cube in a cup.

The hardware bench: an SO-101 arm, a RealSense D455 RGB-D camera on a tripod, and the green bottle that is carried into the arm's path SO-101 arm D455 RGB-D Moving hazard
An SO-101 arm, a RealSense D455 depth camera and a wrist camera, all run by one Intel laptop.

One laptop runs the whole loop

Intel Core Ultra X7 358H

CPU
Safety filter and optical-flow tracker
Integrated GPU
Robot policy, object detector and hazard namer
NPU
Measured 2.5× slower, so unused
What the filter sees Blue shapes cover the arm; the red shape covers the bottle. The filter keeps them apart.
Tracking The red shape moves with the bottle as it is carried; the cyan line shows how far it has gone.

Across four tasks with four tries each, the arm touched the bottle in 3 of 16 tries with the filter and in 11 of 16 without it. It finished the task in 11 tries with the filter and in 13 without.

Small enough for a laptop

Nothing runs on a server or over the network.

~2 ms
for the safety filter to check each command
343 → 177 ms
per policy decision, after trimming unused inputs and taking fewer refinement steps; the policy's weights are unchanged

Abstract

A vision–language–action (VLA) policy can finish a manipulation task while knocking over objects unrelated to it, so task success alone does not show that the policy is safe to deploy in clutter. We study how to keep a pretrained VLA policy clear of such hazards at run time without retraining it, which requires guarding more of the arm than the end effector, following the hazard as it moves, and sharing onboard compute with the policy. Our training-free shield covers the gripper, wrist, and forearm with five ellipsoids and filters every commanded motion through one barrier program against a keep-out ellipsoid fitted from RGB-D perception at reset. Sparse optical flow then carries that ellipsoid's center along with the hazard, with no repeated detection or refitting. Over six simulated hazard-motion conditions, the shield lowers collision from 65.62% to 27.27% and raises safe-success, task completion without collision, from 29.35% to 50.43%. Ablations show that guarding the arm links protects beyond end-effector shielding, and that tracking recovers most of the protection lost when the hazard estimate is frozen at reset. On heterogeneous edge hardware, the five-ellipsoid barrier runs on the CPU in 2.2 ms at the 99th percentile, and trimming the vision–language prefix and taking fewer flow-matching steps shortens each π0.5 policy call on the integrated GPU from 343 to 177.3 ms. On a physical SO-101 arm across four tasks, the arm touched the hazard in 3 of 16 shielded episodes versus 11 of 16 unshielded ones.

BibTeX

@misc{agarwal2026multilink,
  title  = {Multi-Link Safety Filtering for {VLA} Policies
            Around Moving Hazards},
  author = {Agarwal, Yatharth and Raghunathan, Vijay},
  year   = {2026},
  note   = {Preprint}
}