From the 1 of 9 linked papers with an AI index.
9 papers
Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation
Shihao Zhang, Yunzhi Li, Yuguang Yan +4
The paper introduces Shell-LCC, a method that treats the data manifold of high‑quality video training data as an implicit reward model, providing cheap, dense guidance for text‑to‑…
MoCapAnything V2: End-to-End Motion Capture for Arbitrary Skeletons
Kehong Gong, Zhengyu Wen, Dao Thien Phong +10
Recent methods for arbitrary-skeleton motion capture from monocular video follow a factorized pipeline, where a Video-to-Pose network predicts joint positions and an analytical inv…
SpriteHand: Real-Time Versatile Hand-Object Interaction with Autoregressive Video Generation
Zisu Li, Hengye Lyu, Jiaxin Shi +4
Modeling and synthesizing complex hand-object interactions remains a significant challenge, even for state-of-the-art physics engines. Conventional simulation-based approaches rely…
Generative Augmented Reality: Paradigms, Technologies, and Future Applications
Chen Liang, Jiawen Zheng, Yufeng Zeng +7
This paper introduces Generative Augmented Reality (GAR) as a next-generation paradigm that reframes augmentation as a process of world re-synthesis rather than world composition b…
Ca2-VDM: Efficient Autoregressive Video Diffusion Model with Causal Generation and Cache Sharing
Kaifeng Gao, Jiaxin Shi, Hanwang Zhang +3
With the advance of diffusion models, today's video generation has achieved impressive quality. To extend the generation length and facilitate real-world applications, a majority o…
Reasoning Physical Video Generation with Diffusion Timestep Tokens via Reinforcement Learning
Wang Lin, Liyu Jia, Wentao Hu +6
Despite recent progress in video generation, producing videos that adhere to physical laws remains a significant challenge. Traditional diffusion-based methods struggle to extrapol…