activity
20242026
most citedAutoKaggle: A Multi-Agent Framework for Autonomous Data Science Competitions

1 citations · 1 across the 4 of their papers we have counts for

collaborators

6 papers

cs.MA2026

AgentRadio: Passive Awareness for Long-Horizon Multi-Agent Collaboration

Xinxing Ren, Qianbo Zang, Ziyan Wang +4

Understanding large codebases is a long-horizon task for Large Language Model (LLM) agents: answering a single question can require building and running the software, tracing execu…

cs.MA2025

Anemoi: A Semi-Centralized Multi-agent System Based on Agent-to-Agent Communication MCP server from Coral Protocol

Xinxing Ren, Caelum Forder, Qianbo Zang +6

Recent advances in generalist multi-agent systems (MAS) have largely followed a context-engineering plus centralized paradigm, where a planner agent coordinates multiple worker age…

cs.LG2025

SimuGen: Multi-modal Agentic Framework for Constructing Block Diagram-Based Simulation Models

Xinxing Ren, Qianbo Zang, Zekun Guo

Recent advances in large language models (LLMs) have shown impressive performance in mathematical reasoning and code generation. However, LLMs still struggle in the simulation doma…

cs.CV2025

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding

David Ma, Huaqing Yuan, Xingjian Wang +16

Although long-video understanding demands that models capture hierarchical temporal information -- from clip (seconds) and shot (tens of seconds) to event (minutes) and story (hour…

cs.AI20241 cited

AutoKaggle: A Multi-Agent Framework for Autonomous Data Science Competitions

Ziming Li, Qianbo Zang, David Ma +11

Data science tasks involving tabular data present complex challenges that require sophisticated problem-solving approaches. We propose AutoKaggle, a powerful and user-centric frame…

cs.CV2024

LIME: Less Is More for MLLM Evaluation

King Zhu, Qianbo Zang, Shian Jia +18

Multimodal Large Language Models (MLLMs) are evaluated on various benchmarks, such as image captioning, visual question answering, and reasoning. However, many of these benchmarks…