Dex-One2Many

Dex-One2Many: Learning Dexterous Manipulation
from a Single Human Demonstration

1Seoul National University 2University of Maryland, College Park 3KAIST
*Equal contribution †Equal advising

Overview

TL;DR Dex-One2Many learns dexterous manipulation from a single human video and generalizes to poses and grasps never shown in the video. Our key insight is to abstract the video into sequential scene graphs that guide RL, enabling efficient exploration while preserving broad generalizability.

This video includes narration — unmute to listen.

From One Human Video to Many Dexterous Behaviors

A single human video becomes a dexterous policy that generalizes to initial poses, goal poses, and grasps that never appear in the video.

Method

This video includes narration — unmute to listen.

Results

Dex-One2Many Generalizes to Configurations Never Shown in the Video

Seen configurations reproduce the initial and goal poses in the human video; Unseen ones use poses the video never shows. The baselines merely imitate the demonstrated motion, so they succeed only where the scene matches the video. Dex-One2Many goes beyond imitation: it learns the task structure and keeps solving the task when the object starts or ends somewhere the video never showed.

Simulation rollouts on unseen configurations. Five tasks spanning direct manipulation (Doll, Can, Stamp) and tool use (Hammer, Sweep), each learned from one human video and trained for four arm–hand embodiments by swapping only the robot model and its grasp set (94.6–96.8% mean success).

Diverse Grasp Strategies across Dexterous Embodiments and Object Configurations

Rather than retargeting the single grasp shown in the human video, Dex-One2Many trains on diverse grasps synthesized for each hand, so the policy learns grasps that fit its own morphology and the scene. Across the 33-type grasp taxonomy, different embodiments settle on different grasp types, and the same hand changes its grasp as the object configuration changes.

  1. 1Hover a cell to see the graspTap a cell to see the grasp
  2. 2Click to watch the rolloutWatch its rollout

Concurrent Work

Additional works to be added

Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy

Retargets human hand–object interactions across hand morphologies while preserving contacts, refines them with residual RL, and distills the result into zero-shot sim-to-real visuomotor policies.

How it relates

Both learn a dexterous manipulation policy from a single human video. Morphometric Imitation retargets the demonstrated motion, so its generalization is evaluated with the object’s initial position varied only within roughly a 10 cm × 10 cm region around the demonstration. Dex-One2Many abstracts the video into task relations instead of motion, so the policy generalizes to initial and goal poses far from the ones in the video.

FlashDexRetarget: Accelerating Dexterous Manipulation Data Generation through Multi-Motion Retargeting

Learns a single RL policy that tracks many human hand–object demonstrations at once, retargeting them to dexterous hands with about 100× less training compute than per-demonstration methods.

How it relates

Both train a single RL policy that produces diverse dexterous motions. FlashDexRetarget leaves generalization to motions not in the demonstrations as future work. Dex-One2Many targets exactly that case: by abstracting the video into task relations rather than tracking its motion, the policy produces motions that never appear in the video.

DexAgent: An Agentic Human2Sim2Robot Framework for Dexterous Manipulation with Self-Evolving Tool Library

An agentic pipeline that turns one egocentric human video and a task prompt into simulated robot trajectories for policy training, choosing or writing tools for each stage and verifying the outcome.

How it relates

Both augment a single human video into diverse simulated configurations so the robot handles situations the video never shows. They differ in how the robot motion comes about: DexAgent generates trajectories with optimization tools that an agent selects from a self-evolving library and verifies, then trains a policy on that data. Dex-One2Many learns the motion with RL, using the scene graph to vary not only object poses but also the grasp, so each hand discovers its own way to solve the task.

BibTeX

@article{lee2026dexone2many,
  title   = {Dex-One2Many: Learning Dexterous Manipulation from a Single Human Demonstration},
  author  = {Lee, Jusuk and Kim, Sungha and Park, Yeonsoo and Cheon, Jonguk and Jung, Yoonkyo and You, Yongjun and Kim, H. Jin and Huang, Jia-Bin and Huang, Furong and Jang, Youngseok and Lee, Seungjae},
  journal = {arXiv preprint arXiv:2610.12470},
  year    = {2026}
}