From the 1 of 12 linked papers with an AI index.
12 papers
Token Radius Attention for Efficient Video Generation
Jiayu Chen, Zhikun Jiang, Maoliang Li +6
Video Diffusion Transformers (VDiTs) enable high-fidelity generation but incur quadratic cost from dense 3D self-attention. Existing head- and block-level sparse methods share comp…
JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
Shawn Li, Wei Yang, Jike Zhong +11
The paper introduces JigShape, a benchmark of interlocking jigsaw puzzles designed to test visual‑geometric reasoning in vision‑language models, and shows that current zero‑shot an…
Pyramid Forcing: Head-Aware Pyramid KV Cache Policy for High-Quality Long Video Generation
Jiayu Chen, Junbei Tang, Wenbiao Zhao +6
Autoregressive video generation enables streaming and open-ended long video synthesis, but still suffers from long-term degradation caused by accumulated errors. Existing KVCache s…
Representation Fréchet Loss for Visual Generation
Jiawei Yang, Zhengyang Geng, Xuan Ju +2
We show that Fréchet Distance (FD), long considered impractical as a training objective, can in fact be effectively optimized in the representation space. Our idea is simple: deco…
LOME: Learning Human-Object Manipulation with Action-Conditioned Egocentric World Model
Quankai Gao, Jiawei Yang, Qiangeng Xu +2
Learning human-object manipulation presents significant challenges due to its fine-grained and contact-rich nature of the motions involved. Traditional physics-based animation requ…
Fast SAM 3D Body: Accelerating SAM 3D Body for Real-Time Full-Body Human Mesh Recovery
Timing Yang, Sicheng He, Hongyi Jing +4
SAM 3D Body (3DB) achieves state-of-the-art accuracy in monocular 3D human mesh recovery, yet its inference latency of several seconds per image precludes real-time application. We…