HCPG-Flow

Hierarchical Contact-Progress Guidance for Flow-Policy Robot Manipulation

Guanghu Xie, Mingxu Li, Shuo Zhang, Yonglong Zhang, Yifan Yang, Yang Liu, Zongwu Xie, and Baoshi Cao

Demo

Abstract

Flow policies can represent multimodal action distributions for robot manipulation, yet a robot must execute one action at each control step. HCPG-Flow augments SAC-Flow with an analytic, object-centric selector: before contact it favors tool-to-object progress, while after contact it favors object motion along a task-relevant direction. Proposal scores are normalized within the candidate set and converted into a temperature-controlled soft action. The selector introduces no learned parameters or auxiliary backward pass and leaves the original SAC-Flow actor and critic objectives unchanged.

Method

HCPG-Flow framework showing parallel flow proposals, contact-conditioned progress scoring, within-set score normalization, soft action fusion, and the original SAC-Flow learning loop.
Parallel proposals are scored using contact-conditioned geometric progress and fused into one bounded action.

Results

Across-task last-five success improves from 87.2% to 96.7% on ManiSkill and from 94.7% to 97.1% on MetaWorld. Across four physical tasks, HCPG improves pooled success from 91.7% to 98.3% and reduces the macro-average steps to first success from 62.4 to 51.6 (17.4% fewer) relative to SAC-Flow.

Success-rate comparison on four ManiSkill tasks and the across-task average.
ManiSkill. Last-five episode success over seeds 0, 1, and 2.
Success-rate comparison on six MetaWorld tasks and the across-task average.
MetaWorld. Last-five episode success over seeds 0, 1, and 2.

Bars show means; error bars denote standard deviation. They are not confidence intervals or bounded success-rate ranges.

Simulation Videos

ManiSkill

PickCube
PushCube
PokeCube
PullCube

MetaWorld

ButtonPress
DrawerOpen
DoorOpen
SweepInto
PegInsertSide
LeverPull

Real-Robot Videos

Peg insertion
Button press
Drawer closing
Rugby grasp