RASA: Disentangled Spatial-Motional Priors for Cross-Identity Character Animation
Abstract Overview
This paper studies cross-identity character animation, where a reference image provides the target identity and a driving video provides the motion, under challenging mismatches in position, scale, and skeletal proportions. The authors propose RASA, a diffusion-transformer-based framework that explicitly separates spatial alignment from motion control through two modules: the Spatial Prior Calibrator (SPC) and the Inherent Motional Guider (IMG). SPC aligns the driving pose to the reference structure at the start and throughout denoising, while IMG injects shape-agnostic SMPL articulation features into intermediate transformer layers to improve anatomical and view-consistent motion. The paper also introduces CIM-Bench, a manually curated benchmark designed to evaluate realistic cross-identity structural misalignment. Experiments on TikTok and CIM-Bench indicate improved visual quality, identity preservation, and robustness relative to prior animation methods.
Novelty
The main novelty is an explicit disentangling of spatial calibration and motion guidance inside a diffusion transformer for cross-identity animation, rather than treating pose conditioning as a single entangled signal. The work is also distinctive in introducing CIM-Bench, a benchmark focused on realistic structural misalignment across identities.
Results
On TikTok, RASA achieves the best reported image-level metrics in the paper's comparison table, including PSNR 22.03, SSIM 0.830, LPIPS 0.189, and FID 18.93, while remaining competitive on video metrics. On the more challenging CIM-Bench, it outperforms listed baselines across image, video, and identity metrics, for example reaching FVD 667.96, FID-VID 17.63, Sim-Arc 0.72, and Face-FID 90.63. Ablations further show that combining SPC and IMG improves temporal consistency and motion quality beyond either component alone.
Key Points
- RASA uses a dual-prior design: SPC for continuous 2D structural calibration and IMG for identity-agnostic 3D motion semantics.
- The authors build CIM-Bench to evaluate cross-identity animation under realistic pose-reference misalignment, with 656 curated character videos.
- Quantitative and ablation results support that separating spatial alignment from motion control improves robustness, identity consistency, and temporal coherence.