No code

Interspatial Attention for Efficient 4D Human Video Generation

2025 / ACM Transactions on Graphics / DOI 10.1145/3731165

Ruizhi Shao Yinghao Xu Yujun Shen Ceyuan Yang Yang Zheng Changan Chen Yebin Liu Gordon Wetzstein

Generating photorealistic videos of digital humans in a controllable manner is crucial for a plethora of applications. Existing approaches either build on methods that employ template-based 3D representations or emerging video generation models but suffer from poor quality or limited consistency and identity preservation when generating individual or multiple digital humans. In this paper, we introduce a new interspatial attention (ISA) mechanism as a scalable building block for modern diffusion transformer (DiT)-based video generation models. ISA is a new type of cross attention that uses relative positional encodings tailored for the generation of human videos. Leveraging a custom-developed video variation autoencoder, we train a latent ISA-based diffusion model on a large corpus of video data. Our model achieves state-of-the-art performance for 4D human video synthesis, demonstrating remarkable motion consistency and identity preservation while providing precise control of the camera and body poses. Our code and model are publicly released at https://dsaurus.github.io/isa4d/.

Citations

References

Implementations

Reusable

Repro status

Reproducibility Dossier

ReusableConfidence: editor verified / checked Apr 2026

GEOMDIGEST treats reproducibility as an evidence trail: public artifacts, documentation, data, packaging, archival stability, and verification checks. Numeric scores are only exposed for audited records; public pages prioritize the evidence itself.

Evidence

Verified

yes

Code

not yet

Data

not yet

Docs

not yet

Build checks

supplementary / verified / editor verified

Detected evidence link

code / verified / editor verified

Code repository discovered from verified project page

Methodology

Improve this dossier

Implementation Index

No implementations indexed yet

This paper is in the knowledge graph, but we have not attached a runnable artifact yet.

Citation Lineage

References1

2023AvatarReX: Real-time Expressive Full-body Ava...67 cites

Selected paper

Interspatial Attention for Efficient 4D Human Video Gener...

2025 / 1 citations

Cited by0