7 papers
Thinking with Anchors: Grounded and Efficient Document Reasoning
Sichen Zhu, Yuchen Zhu, Wenzhuo Xu +13
Existing document understanding benchmarks have largely focused on locating page elements, yet real-world document intelligence requires models to reason jointly about region seman…
Efficient Reasoning with Hidden Thinking
Xuan Shen, Yizhou Wang, Yufa Zhou +4
Chain-of-Thought (CoT) reasoning has become a powerful framework for improving complex problem-solving capabilities in Multimodal Large Language Models (MLLMs). However, the verbos…
OIDA-QA: A Multimodal Benchmark for Analyzing the Opioid Industry Documents Archive
Xuan Shen, Brian Wingenroth, Zichao Wang +12
The opioid crisis represents a significant moment in public health that reveals systemic shortcomings across regulatory systems, healthcare practices, corporate governance, and pub…
FastCar: Cache Attentive Replay for Fast Auto-Regressive Video Generation on the Edge
Xuan Shen, Weize Ma, Yufa Zhou +11
Auto-regressive (AR) models, initially successful in language generation, have recently shown promise in visual generation tasks due to their superior sampling efficiency. Unlike i…
DraftAttention: Fast Video Diffusion via Low-Resolution Attention Guidance
Xuan Shen, Chenxia Han, Yufa Zhou +7
Diffusion transformer-based video generation models (DiTs) have recently attracted widespread attention for their excellent generation quality. However, their computational cost re…
LazyDiT: Lazy Learning for the Acceleration of Diffusion Transformers
Xuan Shen, Zhao Song, Yufa Zhou +12
Diffusion Transformers have emerged as the preeminent models for a wide array of generative tasks, demonstrating superior performance and efficacy across various applications. The…