SF-Push

Safe and Fair Pose-to-Pose Collaborative Transport
with Multiple Quadruped Robots

Yuhao Li, Williard Joshua Jose, Hao Zhang
Human-Centered Robotics Lab, University of Massachusetts Amherst
Under review at ICRA 2027

94.9%

success on the hardest goals
strongest baseline: 56.7%

88%

of the hardware run with both robots engaged
weighted sum: 52%

0.60 m

minimum inter-robot separation
weighted sum: 0.37 m, inside the 0.5 m limit

This figure shows the baseline, not our method. It is what a hand-weighted scalar reward does on hardware: the box still reaches every waypoint, but the robots close inside the safety distance and one of them stops contributing.
Top-down composite of two Unitree Go1 robots pushing a box through waypoints under a weighted-sum reward, with panels showing inter-robot distance crossing the safety threshold and unbalanced per-robot engagement.

A hand-weighted scalar reward can complete the waypoint task while producing unsafe and unbalanced collaboration on hardware. Left: two Unitree Go1 robots push a box through target poses WP0–WP3 under a weighted sum of task, safety and fairness rewards. Right, top: the inter-robot distance crosses the 0.5 m safety threshold four times, down to 0.37 m. Bottom: during WP0→WP1, R1 is engaged for 96.7% of the segment against 40.8% for R2. The box still reaches every target pose — this is not a team stalled by miscoordination, but a policy that completes the task while one robot contributes little. On the same course SF-Push keeps both robots engaged for 91% and 93% of the run and never comes closer than 0.60 m.

Abstract

Collaborative box pushing by multiple quadrupeds requires the team to reach a target pose, stay safely apart, and keep both robots engaged. Safety and contribution fairness restrict how the team moves without rewarding task completion, so a scalar reward must fix their trade-off with task progress in advance, and with hand-tuned weights both failures appear: one robot disengages for long stretches, and when both push they come closer than the safety distance despite the penalty. We formulate collaborative pushing as a multi-agent multi-objective MDP whose objectives play different roles in the policy update, and train a shared policy with SF-Push, which forms one policy gradient per (robot, objective) pair, takes its direction from a minimum-norm combination of all the unit-normalized gradients, and its scale from the summed task gradient alone, so no trade-off weights are tuned. In simulation SF-Push reaches 99.6% success on nominal goals against 69.1%, 97.3% and 90.6% for a weighted sum, ConFIG and GCR-PPO, with the least time in unsafe proximity of any method on those goals, and keeps 94.9% on the hardest goals, where the strongest baseline reaches 56.7%. On two Unitree Go1 robots SF-Push keeps both engaged for 88% of a three-waypoint course at a minimum separation of 0.60 m, while the weighted-sum policy alternates which robot pushes, keeps both engaged for only 52%, and closes to 0.37 m, inside the 0.5 m requirement.

Simulation results

Collaborative pose-to-pose pushing across three goal settings. Mean ± std over independent training runs. All methods combine the same twelve per-(robot, objective) gradients under identical rewards, critic and hyperparameters, so only the combination rule differs. Bold marks the best value per setting — SF-Push does not lead on every column.

SettingMethodSuccess (%) ↑Time (s) ↓Unsafe prox. (%) ↓Work balance ↑Min. dist. (m) ↑
Easy
0.9 m, 30–60°
Weighted Sum96.39 ± 2.112.63 ± 0.044.230 ± 5.1350.656 ± 0.0290.355 ± 0.053
ConFIG95.14 ± 2.683.44 ± 0.050.128 ± 0.0940.661 ± 0.0540.394 ± 0.081
GCR-PPO94.02 ± 2.962.82 ± 0.431.998 ± 0.5630.649 ± 0.0680.312 ± 0.043
SF-Push98.53 ± 0.764.41 ± 0.520.062 ± 0.0220.660 ± 0.0230.406 ± 0.081
Normal
1.8 m, 90–120°
Weighted Sum69.09 ± 24.275.49 ± 0.773.588 ± 4.9520.771 ± 0.0130.342 ± 0.034
ConFIG97.26 ± 0.305.52 ± 0.370.099 ± 0.0360.803 ± 0.0080.463 ± 0.011
GCR-PPO90.56 ± 8.575.27 ± 0.202.446 ± 0.5570.767 ± 0.0220.322 ± 0.031
SF-Push99.56 ± 0.145.49 ± 0.070.073 ± 0.0250.796 ± 0.0230.457 ± 0.019
Hard
3.6 m, 150–180°
Weighted Sum34.92 ± 25.3311.47 ± 1.552.142 ± 2.0710.750 ± 0.0110.309 ± 0.011
ConFIG40.76 ± 9.559.99 ± 0.280.028 ± 0.0110.731 ± 0.0290.449 ± 0.012
GCR-PPO56.73 ± 34.5811.59 ± 0.930.734 ± 0.2930.746 ± 0.0390.277 ± 0.012
SF-Push94.90 ± 1.699.07 ± 0.770.042 ± 0.0140.806 ± 0.0190.462 ± 0.005

Ablation

Each variant differs from SF-Push in one component: the direction rule, the step-length rule, the gradient basis, or per-robot advantage normalization. Hard setting, same training budget, 3 evaluation seeds. Bold marks the best among variants that reach a nonzero success rate; “—” means the variant never completed the task, so no completion time exists.

VariantSuccess (%) ↑Time (s) ↓Unsafe prox. (%) ↓Work balance ↑Min. dist. (m) ↑
MGDA: raw direction, ‖d*‖ step0.00 ± 0.00—0.000 ± 0.0000.116 ± 0.0200.790 ± 0.029
SF direction, ‖d*‖ step0.00 ± 0.00—0.280 ± 0.0720.695 ± 0.0010.419 ± 0.020
SF direction, unit-norm step65.60 ± 2.7210.26 ± 0.070.194 ± 0.0590.792 ± 0.0080.452 ± 0.010
SF direction, all-row step36.05 ± 1.3814.14 ± 0.280.007 ± 0.0100.739 ± 0.0020.457 ± 0.080
ConFIG direction, task step44.12 ± 1.2711.58 ± 0.120.103 ± 0.0090.742 ± 0.0060.375 ± 0.066
Joint 6-row basis, task step69.14 ± 1.249.72 ± 0.080.041 ± 0.0120.793 ± 0.0090.454 ± 0.018
SF-Push94.57 ± 0.259.54 ± 0.050.039 ± 0.0080.794 ± 0.0080.448 ± 0.025
SF-Push w/o per-robot adv. norm.0.00 ± 0.00—0.501 ± 0.2050.314 ± 0.0200.294 ± 0.022

SF-Push scores 94.57 here and 94.90 in the table above. That is deliberate, not a discrepancy: the ablation is evaluated at a fixed 229,376,000 environment steps over 3 evaluation seeds of a single training run, while the main table uses a larger budget and aggregates over independent training runs.

Method

SF-Push training pipeline: per-(robot, objective) gradients, a scale-free consensus direction, and a step length set by the task gradient.

Six reward terms feed a multi-head critic, and one PPO gradient per (robot, objective) pair gives the rows. Two branches decide different things. Where: every row is unit-normalized and the consensus direction is the minimum-norm point of their convex hull, so it depends only on the angles between rows, not on reward scale. How far: the eight task rows are summed raw. They meet where the task gradient is projected onto the consensus direction to set the step length, and the result enters the unchanged MAPPO optimizer. Only the policy and local observations are needed at deployment. Arrow lengths in stage 1 are log-compressed per-row gradient norms from one representative update; the hull and projection sketches are schematic and carry no numeric claim.

Hardware

Three-waypoint S-curve on two Unitree Go1 robots: engagement and separation lanes for the weighted sum and for SF-Push.

A three-waypoint S-curve spanning 4.45 m in a 3×5 m arena, one trial per policy, with motion capture for poses and no fine-tuning from simulation. Top: weighted sum. Bottom: SF-Push. The four lanes are robot 1 engaged, robot 2 engaged, both engaged, and separation below the 0.5 m threshold. Engaged means the robot's oriented footprint is within 0.10 m of the box — a proximity proxy for participation, as the robots carry no contact-force sensing.

  • The weighted sum spends 2.8% of the run inside the 0.5 m safety distance. SF-Push never enters it.
  • Per-segment engagement under the weighted sum swings hard: R1 96.7% vs R2 40.8% on segment 1, then R2 leads by 28 and 20 points on the two later segments. Under SF-Push the per-segment gap never exceeds 11 points.
  • SF-Push finishes the course in 38.9 s against 55.4 s.

Supplementary video

Time-synced hardware comparison. Left: weighted sum. Right: SF-Push. The current target waypoint is drawn in green with a heading arrow and past waypoints greyed. The “R1 / R2 engaged” badges light up when that robot is within 0.10 m of the box, a red border appears whenever the separation drops below 0.5 m, and the four-lane timeline underneath accumulates as the run proceeds. The running readouts are computed by the same code as the hardware figure above, so they converge to 52% / 0.37 m and 88% / 0.60 m.

Beyond the paper

Material that did not fit the page budget of the submission.

Training environment

Genesis simulation: two Go1 robots at a translucent box on the training floor, with the current box heading in red and the goal heading in green.

Policies are trained and evaluated in Genesis. Two Go1 robots face a box on the training floor; the red arrow is the current box heading and the green arrow the goal heading. Deployment on hardware uses the policy as trained, with no fine-tuning.

Simulation rollouts

Three rendered rollouts on one hard-setting task: weighted sum, ConFIG and SF-Push, each sampled at five points.

The three methods on the same hard-setting task: identical goal pose, 152° of rotation to undo, each row sampled at 0/25/50/75/100% of its own episode. These are Genesis frames, filmed through an orthographic top-down camera in the run that produced the trajectory drawn on them, so the view is a true plan view with no perspective; the green outline is the goal pose with its 0.3 m tolerance and the orange trail is where the box has been. The weighted sum gets the box to the goal but 136° short of the right heading, ConFIG stalls at 0.67 m / 68°, and SF-Push arrives in 10.7 s against the other two timing out at 20 s.

One episode per policy, filmed in the run that produced the trajectory drawn on it. A single pair is an illustration, not evidence: the numbers that are evidence are in the table above.

BibTeX

@misc{li2026sfpush,
  title  = {SF-Push: Safe and Fair Pose-to-Pose Collaborative Transport
            with Multiple Quadruped Robots},
  author = {Li, Yuhao and Jose, Williard Joshua and Zhang, Hao},
  year   = {2026},
  note   = {Under review}
}

Generative AI (Claude, GPT, Kimi K3, and DeepSeek) was used under the authors' direction to draft and revise text in all sections and to write code for training, evaluation, and data processing; all experiments, results, and claims are the authors' own.