10 papers
Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings
Peixi Wu, Ke Mei, Feipeng Ma +15
The paper introduces RIME, a rewrite-driven framework that improves multimodal embeddings by jointly optimizing generation and retrieval-friendly rewriting, aligning generative and…
Preserve, Reveal, Expand: Faithful 4D Video Editing with Region-Aware Conditioning
Zhangchi Hu, Wenzhang Sun, Xiangchen Yin +5
Existing 4D-driven video diffusion models primarily target plausible generation, but faithful 4D editing requires preserving source-observed regions while synthesizing disoccluded…
Q-Probe: Scaling Image Quality Assessment to High Resolution via Context-Aware Agentic Probing
Xiang Li, Xueheng Li, Yu Wang +4
Reinforcement Learning (RL) has empowered Multimodal Large Language Models (MLLMs) to achieve superior human preference alignment in Image Quality Assessment (IQA). However, existi…
Decoupling Perception from Reasoning for Hallucination-Resistant Video Understanding
Bowei Pu, Chuanbin Liu, Yifan Ge +5
Video Large Language Models improve reasoning over complex videos by generating intermediate reasoning text. However, reliable reasoning depends on accurate video perception. In ex…
RiO-DETR: DETR for Real-time Oriented Object Detection
Zhangchi Hu, Yifan Zhao, Yansong Peng +8
We present RiO-DETR: DETR for Real-time Oriented Object Detection, the first real-time oriented detection transformer to the best of our knowledge. Adapting DETR to oriented boundi…
MUSE: A Multi-agent Framework for Unconstrained Story Envisioning via Closed-Loop Cognitive Orchestration
Wenzhang Sun, Zhenyu Wang, Zhangchi Hu +3
Generating long-form audio-visual stories from a short user prompt remains challenging due to an intent-execution gap, where high-level narrative intent must be preserved across co…