SIGGRAPH Asia 2026

UniMate: One Unified Model to Animate Diverse Skeletons

  • Linzhan Mou 1 ·
  • Jiahui Lei 2 ·
  • Zhiyang Dou 3 ·
  • Chenyue Cai 1 ·
  • Chaoyue Song 4 ·
  • Adam Finkelstein 1 ·
  • Szymon Rusinkiewicz 1
  1. 1 Princeton University
  2. 2 University of California, Berkeley
  3. 3 Massachusetts Institute of Technology
  4. 4 Nanyang Technological University

A rigged asset and one sentence are all it takes. UniMate is a single feed-forward model that synthesizes motion for any kinematic tree — bipeds, quadrupeds, birds, fish, insects, and articulated rigid objects — with no template, no per-skeleton fine-tuning, and no reference motion at inference.

viewer

Every clip here is our model's output

Pick a rig, read the prompt it was given, drag to orbit, and scrub the timeline to stop on any pose. Nothing on this page is a video: these are the rigged meshes the model animated, playing in your browser.

research use only

12 rigs · 381 joints

loading dragon-flap.glb…

A dragon flaps its wings.

0 joints · drag to orbit · click a joint

f018/60

Prompts are the demo sheet's own. Several rigs are third-party characters — watch them here, not elsewhere.

one prompt

Same prompt, different bodies

One sentence, three rigs with nothing in common — a stocky armored humanoid, a round soft humanoid, and a quadruped — generated by the same model with no change but the skeleton it is conditioned on. They share a clock, so scrubbing shows where they agree and where the body wins.

one prompt

“A robot walks forward.”

loading 0/3 rigs…

armored humanoid

16 joints

soft humanoid

22 joints

quadruped

19 joints

f018/60

method

One model, any kinematic tree

The model is a flow-matching diffusion transformer. Its one design question is how to put a different graph into the same attention layers, and it answers with three mechanisms — one pairwise, one per joint, one global. Below, the piece that matters: the encoding that replaces a joint's index with its place in the skeleton.

local · pairwise

Graph-aware attention bias

Attention is blind to the kinematic graph. A learned bias, looked up from the pair’s relation type and their hop distance, makes anatomically close joints attend more strongly while long-range coordination stays available — and because the lookup is indexed by relation and distance, never by joint index, it survives an unseen topology.

Bij(h)=(wd(h)) ⁣⊤ed(i,j)+(wr(h)) ⁣⊤er(i,j)B_{ij}^{(h)} = (w_d^{(h)})^{\!\top} e_d(i,j) + (w_r^{(h)})^{\!\top} e_r(i,j)

intrinsic position

Spec-RoPE

Rotary position embedding needs a coordinate for each token. For frames it is the frame index; for joints of a kinematic tree there is no canonical one. Spec-RoPE derives the rotation angle from where the joint sits on the graph instead — the leading eigenvectors of the graph Laplacian — so a joint’s phase is a property of the skeleton, not of how it was serialised.

θj=ω⋅j  ⟶  θj=f(sj)\boldsymbol{\theta}_j = \boldsymbol{\omega} \cdot \htmlClass{text-signal}{j} \;\longrightarrow\; \boldsymbol{\theta}_j = f(\htmlClass{text-signal}{\mathbf{s}_j})

global · per skeleton

Global topological conditioner

A joint-count-invariant summary of the whole skeleton: learnable queries attention-pool the skeleton tokens, and the pooled vector modulates every transformer block through AdaLN-Zero. Local structure comes from the bias, per-joint position from Spec-RoPE, and the overall build of the creature from this.

H=softmax ⁣(1d(Q Wq)(T Wk) ⁣⊤)T Wv\mathbf{H} = \mathrm{softmax}\!\left(\tfrac{1}{\sqrt{d}}(\mathbf{Q}\,W_q)(\mathbf{T}\,W_k)^{\!\top}\right)\mathbf{T}\,W_v

The joint coordinate Spec-RoPE rotates by

the reader has to see it on a rig, not in a formula

building the graph…

Why the index fails

A kinematic tree has no natural order. Serialising it by breadth-first search gives every joint a number, but the numbers say nothing about where a joint sits: index 12 and index 13 can be at opposite ends of the body.

Click any joint, then flip between the two colourings — the index scatters, the spectral coordinate shades smoothly along the body.

Reading the colours

  • BFS index — the order joints are laid out in; colour is position in that list.
  • spectral u₁ — the first non-trivial eigenvector of the graph Laplacian, which Spec-RoPE turns into a rotation angle.
  • graph distance — hops from the joint you clicked, in the rig's own hierarchy.

Each rig is drawn from its own rest pose with the joint hierarchy the model was conditioned on; only the drawing is flat. Eigenvectors are defined up to sign, so a colouring may appear mirrored — the rotations it drives are not.

What the transformer sees, token by token

Every joint contributes one skeleton token — its rest position fused with its parent's — and one motion token per frame, carrying position, a 6D rotation and velocity, plus a depth embedding and a joint-name embedding. Skeleton tokens are prepended along the time axis, so each block attends over structure and motion in one stream.

Attention is factorised: a joint branch across joints at each frame, where the graph bias and Spec-RoPE act, and a temporal branch across frames at each joint with ordinary 1D RoPE. That keeps cost linear in the number of joints instead of quadratic in pairs across all frames.

Skeletons are canonicalised by breadth-first order from the root and normalised by their topology diameter, so scale and serialisation do not leak into the representation.

numbers

What it scores

Two comparisons on held-out skeletons and meshes, and one user study. Held-out here means never seen in training: seven skeleton types, 58 sequences, 5,861 frames for motion generation; 18 meshes across bipeds, quadrupeds, avians, marine life, insects and articulated objects for mesh animation; 32 participants rating 12 cases each for the study.

Motion generation on unseen skeletons

FID between kinematic features of generated and ground-truth motion, and average pairwise joint distance over five samples.

MethodFID ↓Diversity ↑
AnyTop 2.711 8.139
UniMate 0.757 9.2

Mesh animation, against skeleton-free baselines

VBench axes on 512×512 multi-view renders — overall consistency, motion smoothness, dynamic degree, aesthetic quality — and time per animation. AnimateAnyMesh is smoother because its outputs barely move: its dynamic degree is the lowest of the three.

MethodOC ↑MS ↑DD ↑AQ ↑Time ↓
V2M4 0.167 0.991 0.667 0.5061.641 h
AnimateAnyMesh 0.151 0.995 0.352 0.49815.542 s
UniMate 0.186 0.993 0.833 0.5441.214 s

User study

32 participants, 12 cases each, 5-point Likert scale, averaged per method.

MethodText alignmentPlausibilityExpressivenessShape preservationAverage
V2M4 2.378 2.341 2.596 2.336 2.413
AnimateAnyMesh 2.193 2.93 2.362 3.747 2.808
UniMate 4.617 4.568 4.646 4.63 4.615
Ablation: each topology-aware component on its own

Removing the conditioner hurts fidelity most, and its diversity goes up while its motion becomes jittery — the two metrics have to be read together.

VariantFID ↓Diversity ↑
without graph-aware attention bias0.7738.799
without Spec-RoPE0.7988.424
without global topological conditioner0.8259.475
full model0.7579.200

Training data

UniML3D: 13,006 motion sequences (about 20 hours) over thousands of rigs, paired with 3,584 unique text prompts, spanning bipedal, quadrupedal, avian, marine, insectoid, serpentine and articulated-rigid categories — all canonicalised into one representation, with skeletal augmentation applied online during training.

Dataset overview: counts of rigs, sequences and prompts per motion category.

limits

Where it breaks

Two failure modes matter for anyone building on this, and both are structural rather than tuning problems.

Four contact-rich clips showing feet sliding, hovering above the ground and penetrating it.

Contact, footing and drift

There is no unified contact model, because no single one fits every rig here: a foot on the ground is meaningful for a biped or a quadruped, and ill-defined for a snake, a swimming fish, a bird in flight or an articulated lamp. Contact-rich locomotion can therefore slide, drift, hover or sink. Foot locking or IK on the predicted rig fixes a good deal of it where contacts are well defined, and a morphology-aware contact term during sampling is the obvious next step.

Generated motion on rare or unusual skeletons that comes out static, jittery or semantically wrong.

Rare topologies and rare motions

The training data is long-tailed and leans humanoid, so unusual skeletons and out-of-distribution prompts are where outputs become static, jittery or simply wrong about the verb. Broadening coverage — video priors, or scripted and simulated motion over generated assets — is the direction we would take, and nothing about the architecture blocks it.

cite

Cite

UniMate: One Unified Model to Animate Diverse Skeletons — SIGGRAPH Asia 2026 Conference Papers.

BibTeX

@inproceedings{mou2026unimate,
  author    = {Mou, Linzhan and Lei, Jiahui and Dou, Zhiyang and Cai, Chenyue and
               Song, Chaoyue and Finkelstein, Adam and Rusinkiewicz, Szymon},
  title     = {{UniMate}: One Unified Model to Animate Diverse Skeletons},
  booktitle = {SIGGRAPH Asia 2026 Conference Papers},
  series    = {SA Conference Papers '26},
  year      = {2026},
  month     = dec,
  location  = {Kuala Lumpur, Malaysia},
  publisher = {ACM},
  doi       = {10.1145/3829340.3842216},
  isbn      = {979-8-4007-2842-6/2026/12}
}

UniMate · SIGGRAPH Asia 2026 Conference Papers · DOI 10.1145/3829340.3842216