From the 1 of 15 linked papers with an AI index.
1 citations · 1 across the 12 of their papers we have counts for
7 papers · 1 filter
StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding
Xichen Zhang, Guankai Li, Yinghao Zhu +6
Deploying autonomous multimodal agents in continuous, real-world environments requires them to ingest unbounded audio-visual streams and maintain hour-scale memory. However, curren…
VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation
Kangning Zhang, Yixing Li, Shuai Shao +9
The paper proposes Visual Attribution Distillation (VAD), a counterfactual method that isolates the visual component of teacher corrections in multimodal on‑policy distillation and…
MuSEAgent: A Multimodal Reasoning Agent with Stateful Experiences
Shijian Wang, Jiarui Jin, Runhao Fu +11
Research agents have recently achieved significant progress in information seeking and synthesis across heterogeneous textual and visual sources. In this paper, we introduce MuSEAg…
GlyphBanana: Advancing Precise Text Rendering Through Agentic Workflows
Zexuan Yan, Jiarui Jin, Yue Ma +5
Despite recent advances in generative models driving significant progress in text rendering, accurately generating complex text and mathematical formulas remains a formidable chall…
Synthetic Curriculum Reinforces Compositional Text-to-Image Generation
Shijian Wang, Runhao Fu, Siyi Zhao +6
Text-to-Image (T2I) generation has long been an open problem, with compositional synthesis remaining particularly challenging. This task requires accurate rendering of complex scen…
Video-Thinker: Sparking "Thinking with Videos" via Reinforcement Learning
Shijian Wang, Jiarui Jin, Xingjian Wang +6
Recent advances in image reasoning methods, particularly "Thinking with Images", have demonstrated remarkable success in Multimodal Large Language Models (MLLMs); however, this dyn…