25 papers
OmniCoT: A Benchmark for Global and Multi-Step Panoramic Reasoning
Haocong He, Chenfei Liao, Zichen Wen +13
Multimodal Large Language Models (MLLMs) have demonstrated promising spatial reasoning capabilities, while these abilities remain underexplored in the emerging visual modality of p…
Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression Methods
Chenfei Liao, Wensong Wang, Zichen Wen +10
Recent efforts to accelerate inference in Multimodal Large Language Models (MLLMs) have largely focused on visual token compression. The effectiveness of these methods is commonly…
StruVis: Enhancing Reasoning-based Text-to-Image Generation via Thinking with Structured Vision
Yuanhuiyi Lyu, Kaiyu Lei, Ziqiao Weng +7
Reasoning-based text-to-image (T2I) generation requires models to interpret complex prompts accurately. Existing reasoning frameworks can be broadly categorized into two types: (1)…
Rethinking Token-wise Feature Caching: Accelerating Diffusion Transformers with Dual Feature Caching
Chang Zou, Evelyn Zhang, Shikang Zheng +6
Diffusion Transformers (DiT) have become the dominant methods in image and video generation yet still suffer substantial computational costs. As an effective approach for DiT accel…
Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks
Xu Zheng, Zihao Dongfang, Lutao Jiang +17
Humans possess spatial reasoning abilities that enable them to understand spaces through multimodal observations, such as vision and sound. Large multimodal reasoning models extend…
EEdit: Rethinking the Spatial and Temporal Redundancy for Efficient Image Editing
Zexuan Yan, Yue Ma, Chang Zou +3
Inversion-based image editing is rapidly gaining momentum while suffering from significant computation overhead, hindering its application in real-time interactive scenarios. In th…