A rigged asset and one sentence are all it takes. UniMate is a single feed-forward model that synthesizes motion for any kinematic tree — bipeds, quadrupeds, birds, fish, insects, and articulated rigid objects — with no template, no per-skeleton fine-tuning, and no reference motion at inference.
Pick a rig, read the prompt it was given, drag to orbit, and scrub the timeline to stop on any pose. Nothing on this page is a video: these are the rigged meshes the model animated, playing in your browser.
research use only
12 rigs · 381 joints
A dragon flaps its wings.
0 joints · drag to orbit · click a joint
Prompts are the demo sheet's own. Several rigs are third-party characters — watch them here, not elsewhere.
One sentence, three rigs with nothing in common — a stocky armored humanoid, a round soft humanoid, and a quadruped — generated by the same model with no change but the skeleton it is conditioned on. They share a clock, so scrubbing shows where they agree and where the body wins.
one prompt
“A robot walks forward.”
armored humanoid
16 joints
soft humanoid
22 joints
quadruped
19 joints
The model is a flow-matching diffusion transformer. Its one design question is how to put a different graph into the same attention layers, and it answers with three mechanisms — one pairwise, one per joint, one global. Below, the piece that matters: the encoding that replaces a joint's index with its place in the skeleton.
local · pairwise
Attention is blind to the kinematic graph. A learned bias, looked up from the pair’s relation type and their hop distance, makes anatomically close joints attend more strongly while long-range coordination stays available — and because the lookup is indexed by relation and distance, never by joint index, it survives an unseen topology.
intrinsic position
Rotary position embedding needs a coordinate for each token. For frames it is the frame index; for joints of a kinematic tree there is no canonical one. Spec-RoPE derives the rotation angle from where the joint sits on the graph instead — the leading eigenvectors of the graph Laplacian — so a joint’s phase is a property of the skeleton, not of how it was serialised.
global · per skeleton
A joint-count-invariant summary of the whole skeleton: learnable queries attention-pool the skeleton tokens, and the pooled vector modulates every transformer block through AdaLN-Zero. Local structure comes from the bias, per-joint position from Spec-RoPE, and the overall build of the creature from this.
the reader has to see it on a rig, not in a formula
A kinematic tree has no natural order. Serialising it by breadth-first search gives every joint a number, but the numbers say nothing about where a joint sits: index 12 and index 13 can be at opposite ends of the body.
Click any joint, then flip between the two colourings — the index scatters, the spectral coordinate shades smoothly along the body.
Each rig is drawn from its own rest pose with the joint hierarchy the model was conditioned on; only the drawing is flat. Eigenvectors are defined up to sign, so a colouring may appear mirrored — the rotations it drives are not.
Every joint contributes one skeleton token — its rest position fused with its parent's — and one motion token per frame, carrying position, a 6D rotation and velocity, plus a depth embedding and a joint-name embedding. Skeleton tokens are prepended along the time axis, so each block attends over structure and motion in one stream.
Attention is factorised: a joint branch across joints at each frame, where the graph bias and Spec-RoPE act, and a temporal branch across frames at each joint with ordinary 1D RoPE. That keeps cost linear in the number of joints instead of quadratic in pairs across all frames.
Skeletons are canonicalised by breadth-first order from the root and normalised by their topology diameter, so scale and serialisation do not leak into the representation.
Two comparisons on held-out skeletons and meshes, and one user study. Held-out here means never seen in training: seven skeleton types, 58 sequences, 5,861 frames for motion generation; 18 meshes across bipeds, quadrupeds, avians, marine life, insects and articulated objects for mesh animation; 32 participants rating 12 cases each for the study.
FID between kinematic features of generated and ground-truth motion, and average pairwise joint distance over five samples.
| Method | FID ↓ | Diversity ↑ |
|---|---|---|
| AnyTop | 2.711 | 8.139 |
| UniMate | 0.757 | 9.2 |
VBench axes on 512×512 multi-view renders — overall consistency, motion smoothness, dynamic degree, aesthetic quality — and time per animation. AnimateAnyMesh is smoother because its outputs barely move: its dynamic degree is the lowest of the three.
| Method | OC ↑ | MS ↑ | DD ↑ | AQ ↑ | Time ↓ |
|---|---|---|---|---|---|
| V2M4 | 0.167 | 0.991 | 0.667 | 0.506 | 1.641 h |
| AnimateAnyMesh | 0.151 | 0.995 | 0.352 | 0.498 | 15.542 s |
| UniMate | 0.186 | 0.993 | 0.833 | 0.544 | 1.214 s |
32 participants, 12 cases each, 5-point Likert scale, averaged per method.
| Method | Text alignment | Plausibility | Expressiveness | Shape preservation | Average |
|---|---|---|---|---|---|
| V2M4 | 2.378 | 2.341 | 2.596 | 2.336 | 2.413 |
| AnimateAnyMesh | 2.193 | 2.93 | 2.362 | 3.747 | 2.808 |
| UniMate | 4.617 | 4.568 | 4.646 | 4.63 | 4.615 |
Removing the conditioner hurts fidelity most, and its diversity goes up while its motion becomes jittery — the two metrics have to be read together.
| Variant | FID ↓ | Diversity ↑ |
|---|---|---|
| without graph-aware attention bias | 0.773 | 8.799 |
| without Spec-RoPE | 0.798 | 8.424 |
| without global topological conditioner | 0.825 | 9.475 |
| full model | 0.757 | 9.200 |
UniML3D: 13,006 motion sequences (about 20 hours) over thousands of rigs, paired with 3,584 unique text prompts, spanning bipedal, quadrupedal, avian, marine, insectoid, serpentine and articulated-rigid categories — all canonicalised into one representation, with skeletal augmentation applied online during training.

Two failure modes matter for anyone building on this, and both are structural rather than tuning problems.

There is no unified contact model, because no single one fits every rig here: a foot on the ground is meaningful for a biped or a quadruped, and ill-defined for a snake, a swimming fish, a bird in flight or an articulated lamp. Contact-rich locomotion can therefore slide, drift, hover or sink. Foot locking or IK on the predicted rig fixes a good deal of it where contacts are well defined, and a morphology-aware contact term during sampling is the obvious next step.

The training data is long-tailed and leans humanoid, so unusual skeletons and out-of-distribution prompts are where outputs become static, jittery or simply wrong about the verb. Broadening coverage — video priors, or scripted and simulated motion over generated assets — is the direction we would take, and nothing about the architecture blocks it.
UniMate: One Unified Model to Animate Diverse Skeletons — SIGGRAPH Asia 2026 Conference Papers.
@inproceedings{mou2026unimate,
author = {Mou, Linzhan and Lei, Jiahui and Dou, Zhiyang and Cai, Chenyue and
Song, Chaoyue and Finkelstein, Adam and Rusinkiewicz, Szymon},
title = {{UniMate}: One Unified Model to Animate Diverse Skeletons},
booktitle = {SIGGRAPH Asia 2026 Conference Papers},
series = {SA Conference Papers '26},
year = {2026},
month = dec,
location = {Kuala Lumpur, Malaysia},
publisher = {ACM},
doi = {10.1145/3829340.3842216},
isbn = {979-8-4007-2842-6/2026/12}
}