FuguReport

RASA: Disentangled Spatial-Motional Priors for Cross-Identity Character Animation

Authors Zhen Xiao, Zhen Shen, Zhaofan Qiu, Ting Yao, Xueliang Liu, Tao Mei
Affiliations HiDream.ai Inc. / Hefei University of Technology
Categories Method / Motion Control / Framework separating motion and spatial mapping, Application / Character Animation / Cross-identity driven animation, Research / Animation Robustness / Handling spatial-motion inconsistency
License CC BY 4.0

Abstract Overview

This paper studies cross-identity character animation, where a reference image provides the target identity and a driving video provides the motion, under challenging mismatches in position, scale, and skeletal proportions. The authors propose RASA, a diffusion-transformer-based framework that explicitly separates spatial alignment from motion control through two modules: the Spatial Prior Calibrator (SPC) and the Inherent Motional Guider (IMG). SPC aligns the driving pose to the reference structure at the start and throughout denoising, while IMG injects shape-agnostic SMPL articulation features into intermediate transformer layers to improve anatomical and view-consistent motion. The paper also introduces CIM-Bench, a manually curated benchmark designed to evaluate realistic cross-identity structural misalignment. Experiments on TikTok and CIM-Bench indicate improved visual quality, identity preservation, and robustness relative to prior animation methods.

Novelty

The main novelty is an explicit disentangling of spatial calibration and motion guidance inside a diffusion transformer for cross-identity animation, rather than treating pose conditioning as a single entangled signal. The work is also distinctive in introducing CIM-Bench, a benchmark focused on realistic structural misalignment across identities.

Results

On TikTok, RASA achieves the best reported image-level metrics in the paper's comparison table, including PSNR 22.03, SSIM 0.830, LPIPS 0.189, and FID 18.93, while remaining competitive on video metrics. On the more challenging CIM-Bench, it outperforms listed baselines across image, video, and identity metrics, for example reaching FVD 667.96, FID-VID 17.63, Sim-Arc 0.72, and Face-FID 90.63. Ablations further show that combining SPC and IMG improves temporal consistency and motion quality beyond either component alone.

Key Points

  1. RASA uses a dual-prior design: SPC for continuous 2D structural calibration and IMG for identity-agnostic 3D motion semantics.
  2. The authors build CIM-Bench to evaluate cross-identity animation under realistic pose-reference misalignment, with 656 curated character videos.
  3. Quantitative and ablation results support that separating spatial alignment from motion control improves robustness, identity consistency, and temporal coherence.

References

This page was created using generative AI such as GPT-5, Claude Opus 4, Gemini 3, Gemini 3.1 Flash Image, and their higher-end successor versions. No guarantee can be made regarding its contents.