BVH  ·  SMPL  ·  AMASS  →  URDF, MJCF

Motion retargeting, from a person to a robot

Open a motion capture file, open a robot, and the robot moves the way the person did — as nearly as its joints, limits and proportions allow. Then the part nobody else answers: run the result through physics and find out whether the motors in the robot's own description can actually perform it.

What retargeting has to solve

A recording says how a person's bones turned. A robot has different bones: other lengths, fewer or more joints, axes laid out by an engineer rather than by anatomy, limits a person does not have and a rest pose that is not a person's. There is no correct transfer between the two — only defensible ones — and the differences show up in the same places every time: a shoulder that cannot reach where an arm went, a knee asked to bend past its stop, a foot that slides because the robot's leg is shorter than the step.

So retargeting is two questions, and most tools answer only the first. What should each joint do to look like the recording? And can this machine do it — with these motors, this mass, on its own two feet?

How it works here

Humanoids: matching where the bones point

For a robot with two legs and two arms, every frame of the recording is measured in the frame of the person's pelvis — so walking across a room and turning on the spot do not leak into the joints — and turned into directions: where the thigh points, the shin, the upper arm, the forearm, and how the shoulders sit over the hips. Each limb of the robot is then solved for the joint angles that point its own bones the same way, within its own limits, starting from the solution a frame earlier so a limb does not flip between two equally good answers mid-stride.

Nothing about the recording's conventions is assumed. Left is wherever the left hip is, up is wherever the shoulders are, so a rig that faces the other way or stands Z-up works the same. On the robot side the bones are measured between joint centres, not joint origins — a hip whose three axes do not quite meet is measured from where they come closest, which is what keeps a straight human leg a straight robot leg.

Everything else: one channel per joint

A quadruped shares no body plan with a person, and pretending otherwise produces nonsense with confidence. There each robot joint is driven by one rotation channel of one recorded bone — front legs from the arms, rear legs from the legs — scaled to fit the joint's own range and centred on the robot's stance. The guess is shown, and every row can be changed: the mapping is something you can read and argue with, which an optimiser's output is not.

What it reads

  • BVH, from any rig whose bones are named like a person's — CMU, Mixamo, most exporters. The recorded person can be drawn beside the robot to compare by eye.
  • SMPL, SMPL-H and SMPL-X motion as .npz, the layout AMASS uses: the 22 joints of the body are read, at the file's own frame rate; fingers, jaw and eyes are left alone.
  • Any robot the site opens — URDF, Xacro, MJCF, SDF, USD — your own upload or one of the hundred in the catalogue.

The result is a clip on the timeline like any other: keyframes you can edit, thinned from the recording's rate to what a person can work with, and exportable as a ROS JointTrajectory, for the Unitree SDK, or as CSV and JSON.

Then the part nobody else answers

Press Check and the clip is played through MuJoCo with the robot's own masses and the effort its description declares for each joint. What comes back is specific: which joint is asked for more torque than its motor has, and when; which one rests on its limit; where a foot slides; whether the robot is still upright at the end. A human walk retargeted onto a robot often fails here for a reason no amount of looking at the animation would show — and finding that out before the robot does is the point.

AMASS, SMPL and SMPL-X

SMPL is a parametric model of the human body: a fixed skeleton of joints, each frame of a motion stored as one rotation per joint. SMPL-H adds the hands and SMPL-X the hands and the face. AMASS is a large collection of motion capture datasets converted into those bodies, which is why it is where most humanoid motion work starts. Its files are distributed after registration, under its own licence; this tool reads the ones you have, in your tab.

Where it stops being the right tool

  • Feet are not pinned to the ground while retargeting — a robot with shorter legs will slide, and Check says where.
  • The person's travel through the room is not replayed; balance is the robot's to find.
  • Hands, fingers and the face are not mapped.
  • It is one clip on one robot, interactively. For preparing a dataset to train a policy, an offline pipeline such as GMR — General Motion Retargeting — is built for that job.

Common questions

What is motion retargeting?

Taking a movement recorded on one body and replaying it on another with different proportions, joints and limits — here, from a person in a motion capture file to a robot described in URDF or MJCF. The answer is never exact, because the two bodies are not the same; the useful question is whether the robot can perform the result.

Which files can I retarget?

BVH from any rig whose bones are named like a person's — CMU, Mixamo and most exporters — and SMPL, SMPL-H or SMPL-X motion saved as .npz, the way AMASS distributes it. The body's 22 joints are used; hands and face are left alone.

Can I retarget AMASS to a Unitree G1 or H1?

Yes. Open the robot from the catalogue in Motion, switch to the Capture tab and open the .npz. Legs, arms and waist follow the recording; Check then runs the result through MuJoCo against the motors the description declares.

Does the robot travel the way the person did?

No — joint angles are retargeted, the path through the room is not. A person's balance is not a robot's, and replaying the pelvis's travel would hide exactly what the physics check is there to find: whether the robot stays upright when its own feet have to carry it.

How is this different from GMR?

General Motion Retargeting is an open Python project that solves inverse kinematics on every frame, with scaling between the human and the robot, to produce reference motion for training humanoid policies offline. This runs in a tab, needs no install and ends with a physics check. For a large dataset feeding a training pipeline, use GMR; for looking at one clip on one robot and asking whether its motors can do it, this is faster.

Why does the retargeted robot fall over?

Because a person keeps their balance with a body the robot does not have — different mass, feet and reach. The retarget says what each joint should do; the physics lab and the Check in Motion say what happens when the robot tries, and which joint gave up first.

Where to go next