← Back to feed
2026-06-30visionmultimodal

LUNA: Learning Universal 3D Human Animation Beyond Skinning

Peng Li, Rawal Khirodkar, Junxuan Li, Yuan Dong, Chen Cao, Yuan Liu, Wenhan Luo, Yike Guo, Shunsuke Saito

PDF preview unavailable
Read on arXiv →

Key claim

LUNA enables realistic 3D animation from 2D inputs without fitting.

In plain English

Imagine you want to create lifelike 3D avatars of people just from simple 2D images. Traditionally, this involves complex models that fit a person's body to a predefined shape, which can lead to awkward movements and visual artifacts. This is especially problematic when the input images vary widely or when you want to animate different characters without starting from scratch. These issues are known as fitting artifacts and expressivity constraints.

What LUNA does is quite different. Instead of relying on these traditional fitting methods, it uses a neural network to directly translate various 2D inputs—like images, sketches, or keypoints—into 3D movements. The core of this approach is a transformer-based model that separates the overall motion from the finer details, allowing it to capture both broad movements and subtle nuances in how a character moves. This means you can create more expressive and realistic animations without the usual constraints.

The results are promising: LUNA not only matches the visual quality of existing methods but also allows for realistic animations across different characters and styles without needing extensive retraining. This could be a game-changer for anyone looking to build applications in gaming, film, or virtual reality where realistic human motion is crucial.

Novelty
8.0/10

LUNA introduces a novel approach to 3D human animation without relying on traditional body fitting methods.

Reliability
7.5/10

The claims are supported by extensive experiments demonstrating competitive visual fidelity and generalization.

LUNA: Learning Universal 3D Human Animation Beyond Skinning — Frontier Papers