works on

From the 1 of 15 linked papers with an AI index.

most citedOmniGAIA: Towards Native Omni-Modal AI Agents

1 citations · 1 across the 12 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding

Xichen Zhang, Guankai Li, Yinghao Zhu +6

Deploying autonomous multimodal agents in continuous, real-world environments requires them to ingest unbounded audio-visual streams and maintain hour-scale memory. However, curren…

cs.CV2026

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation

Kangning Zhang, Yixing Li, Shuai Shao +9

The paper proposes Visual Attribution Distillation (VAD), a counterfactual method that isolates the visual component of teacher corrections in multimodal on‑policy distillation and…

cs.CV2026

MuSEAgent: A Multimodal Reasoning Agent with Stateful Experiences

Shijian Wang, Jiarui Jin, Runhao Fu +11

Research agents have recently achieved significant progress in information seeking and synthesis across heterogeneous textual and visual sources. In this paper, we introduce MuSEAg…

cs.CV2026

GlyphBanana: Advancing Precise Text Rendering Through Agentic Workflows

Zexuan Yan, Jiarui Jin, Yue Ma +5

Despite recent advances in generative models driving significant progress in text rendering, accurately generating complex text and mathematical formulas remains a formidable chall…

cs.CV2025

Synthetic Curriculum Reinforces Compositional Text-to-Image Generation

Shijian Wang, Runhao Fu, Siyi Zhao +6

Text-to-Image (T2I) generation has long been an open problem, with compositional synthesis remaining particularly challenging. This task requires accurate rendering of complex scen…

cs.CV2025

Video-Thinker: Sparking "Thinking with Videos" via Reinforcement Learning

Shijian Wang, Jiarui Jin, Xingjian Wang +6

Recent advances in image reasoning methods, particularly "Thinking with Images", have demonstrated remarkable success in Multimodal Large Language Models (MLLMs); however, this dyn…