works on

From the 1 of 18 linked papers with an AI index.

most citedOmniGAIA: Towards Native Omni-Modal AI Agents

1 citations · 1 across the 12 of their papers we have counts for

collaborators
Showing cs.CVShow all

9 papers · 1 filter

cs.CV2026

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding

Xichen Zhang, Guankai Li, Yinghao Zhu +6

Deploying autonomous multimodal agents in continuous, real-world environments requires them to ingest unbounded audio-visual streams and maintain hour-scale memory. However, curren…

cs.CV2026

EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE

Zexuan Yan, Yuzhou Wu, Yue Ma +9

The paper introduces EgoGenesis, a simulator that generates controllable egocentric manipulation videos using geometry-aware conditioning mechanisms to augment real robot data and…

cs.CV2026

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

Fanqing Meng, Lingxiao Du, Zijian Wu +46

Language-model agents are increasingly used as persistent coworkers that assist users across multiple working days. During such workflows, the surrounding environment may change in…

cs.CV2026

MuSEAgent: A Multimodal Reasoning Agent with Stateful Experiences

Shijian Wang, Jiarui Jin, Runhao Fu +11

Research agents have recently achieved significant progress in information seeking and synthesis across heterogeneous textual and visual sources. In this paper, we introduce MuSEAg…

cs.CV2026

GlyphBanana: Advancing Precise Text Rendering Through Agentic Workflows

Zexuan Yan, Jiarui Jin, Yue Ma +5

Despite recent advances in generative models driving significant progress in text rendering, accurately generating complex text and mathematical formulas remains a formidable chall…

cs.CV2025

Synthetic Curriculum Reinforces Compositional Text-to-Image Generation

Shijian Wang, Runhao Fu, Siyi Zhao +6

Text-to-Image (T2I) generation has long been an open problem, with compositional synthesis remaining particularly challenging. This task requires accurate rendering of complex scen…