From the 1 of 8 linked papers with an AI index.
8 papers
TopoAgent: A Self-Evolving Topological Agent for Multimodal Scientific Reasoning
Mingze Xu, Yinghui Li, Jiayi Kuang +5
TopoAgent introduces a graph‑based, self‑evolving framework that breaks down multimodal scientific queries into visual atoms and organizes them in a DAG, allowing dynamic refinemen…
ANX: Protocol-First Design for AI Agent Interaction with a Supporting 3EX Decoupled Architecture
Xu Mingze
AI agents, autonomous digital actors, need agent-native protocols; existing methods include GUI automation and MCP-based skills, with defects of high token consumption, fragmented…
VideoCoF: Unified Video Editing with Temporal Reasoner
Xiangpeng Yang, Ji Xie, Yiyuan Yang +4
Existing video editing methods face a critical trade-off: expert models offer precision but rely on task-specific priors like masks, hindering unification; conversely, unified temp…
AToken: A Unified Tokenizer for Vision
Jiasen Lu, Liangchen Song, Mingze Xu +5
We present AToken, the first unified visual tokenizer that achieves both high-fidelity reconstruction and semantic understanding across images, videos, and 3D assets. Unlike existi…
StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming Assistant
Haibo Wang, Bo Feng, Zhengfeng Lai +6
We present StreamBridge, a simple yet effective framework that seamlessly transforms offline Video-LLMs into streaming-capable models. It addresses two fundamental challenges in ad…
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
Rui Tian, Mingfei Gao, Mingze Xu +5
We introduce UniGen, a unified multimodal large language model (MLLM) capable of image understanding and generation. We study the full training pipeline of UniGen from a data-centr…