End-to-end pipeline that takes a single 2D character image and produces a rigged, animated 3D GLB. Built on Modal for serverless GPU inference.
input image --> InstantMesh (3D mesh)
|
RigAnything (auto-rigging) + motion generation (AnyTop or MDM)
| |
+--------- retarget -----------+
|
animated GLB
| Stage | Model | GPU | Description |
|---|---|---|---|
| Mesh generation | InstantMesh + Zero123++ | A10G | Multi-view reconstruction from a single image |
| Auto-rigging | RigAnything | A10G | Predicts skeleton and skinning weights |
| Motion synthesis | AnyTop or MDM (see below) | A10G | Generates motion sequences |
| Retargeting | Custom FK/IK solver | CPU | Transfers motion onto the rigged character mesh |
Stages 2 (rigging) and 3 (motion) run in parallel via modal.spawn(). Typical end-to-end latency: ~45s warm, ~50s cold.
The pipeline implements two interchangeable motion generation backends, each with distinct trade-offs.
Gat, Raab, Tevet, Atzmon, Chechik, Cohen-Or. "AnyTop: Open-Vocabulary Topology-Agnostic Motion Generation", 2025. [Paper] [Code] [Weights]
AnyTop is a diffusion model that conditions on skeleton topology (parent chain, bone offsets, T5-encoded joint names), enabling motion generation for arbitrary skeletons — humanoids, quadrupeds, or any rigged character.
Strengths:
- Works with any skeleton topology, not limited to humanoids
- Generates motion directly for the target skeleton — no cross-skeleton retargeting needed
- Leverages T5 text embeddings of joint names for semantic understanding of novel skeletons
Limitations:
- Does not support free-form text prompts for motion description (conditions on skeleton semantics, not action descriptions)
- Relatively new; quality on standard humanoid benchmarks may lag behind SMPL-specialized models
Status: Fully integrated and working end-to-end. Since AnyTop generates motion in the character's own skeleton space, the retargeting is a direct joint-name lookup with no topology mismatch.
Tevet, Raab, Gordon, Shafir, Cohen-Or, Bermano. "Human Motion Diffusion Model", ICLR 2023. [Paper] [Code]
MDM is a classifier-free diffusion model that generates motion in HumanML3D representation (263-dim) for a fixed SMPL-22 skeleton. It uses a transformer encoder with CLIP text conditioning, predicting clean samples (x_0) at each diffusion step.
Strengths:
- Supports text-to-motion: describe any action in natural language ("a person walks forward and waves")
- Mature and well-established (ICLR 2023), strong results on HumanML3D benchmarks
- Simple architecture (no VAE/latent space), fast 50-step DDPM sampling
Limitations:
- Fixed to SMPL-22 skeleton — cannot handle non-humanoid or non-standard topologies
- Requires an additional skeleton retargeting step: MDM outputs SMPL-22 joint positions, which must be converted to the character's skeleton via SMPLify (position-to-rotation fitting) and analytical IK
Status: Integrated with SMPLify-based retargeting (motion_mdm.py). The MDM-to-character retargeting involves fitting SMPL body model rotations to MDM's output positions, then transferring those rotations onto the character rig. Due to the skeleton mismatch between SMPL's canonical bone proportions and the character's actual proportions, the resulting animation exhibits pose artifacts — particularly twist accumulation in limbs and imprecise joint angles. This is a fundamental challenge of cross-skeleton retargeting when bone lengths and rest poses differ.
| AnyTop | MDM | |
|---|---|---|
| Text-to-motion | No | Yes |
| Skeleton support | Any topology | SMPL-22 only |
| Retarget complexity | Direct (same skeleton) | SMPLify + IK (cross-skeleton) |
| Animation quality | Clean (no retarget artifacts) | Pose artifacts from skeleton mismatch |
| Best for | Non-humanoid / custom rigs | Text-driven humanoid motion |
character-pipeline/ # Core pipeline (Modal deployment)
stages/
mesh_gen.py # Stage 1: image -> 3D mesh
rigging.py # Stage 2: mesh -> rigged skeleton
motion.py # Stage 3a: motion generation (AnyTop)
motion_mdm.py # Stage 3b: motion generation (MDM + SMPLify)
retarget.py # Stage 4: motion + rig -> animated GLB
api.py # FastAPI REST endpoints
cli.py # CLI interface
config.py # Pydantic config & request/response schemas
app.py # Modal app entrypoint
scripts/ # Standalone utilities
collapse_skeleton_smpl22.py # Convert RigAnything skeleton to SMPL-22 topology
mdm_motion_studio.py # Web UI for text-to-motion visualization (three.js)
setup_mdm_weights.py # Upload MDM dataset stats to Modal volume
setup_smplify_weights.py # Upload SMPL body model files to Modal volume
inspect_glb_animation.py # Inspect animation channels in a GLB file
# Install dependencies
curl -LsSf https://astral.sh/uv/install.sh | sh
cd character-pipeline && uv sync
# Configure Modal
pip install modal && modal setup
# (Optional) HuggingFace token for gated models
modal secret create huggingface HF_TOKEN=hf_...
# Deploy to Modal (first deploy downloads weights, ~10 min)
uv run modal deploy character_pipeline/app.pyFor MDM, additional setup is required:
# Upload SMPL body model and GMM prior to Modal volume
python scripts/setup_smplify_weights.py
python scripts/setup_mdm_weights.py# Health check
curl https://<modal-url>/health
# Generate animated character (multipart upload)
curl -X POST https://<modal-url>/generate/upload \
-F "image=@character.png" \
--output character.glbuv run python -m character_pipeline.cli generate \
--image character.png \
--output character.glbInteractive three.js viewer for previewing MDM text-to-motion results:
python scripts/mdm_motion_studio.py --rig path/to/rigged.glb
# Opens browser at http://localhost:8421| Mechanism | Effect |
|---|---|
@modal.enter(snap=True) — memory snapshot after model load |
Cold start: 45s -> ~2s |
keep_warm=1 — always 1 hot GPU container |
Zero cold starts under load |
modal.Dict cache — SHA-256 keyed result store |
Identical inputs: <1ms |
spawn() + gather() — rig and motion run in parallel |
Saves ~20s per request |
MIT