7 papers
StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding
Xichen Zhang, Guankai Li, Yinghao Zhu +6
Deploying autonomous multimodal agents in continuous, real-world environments requires them to ingest unbounded audio-visual streams and maintain hour-scale memory. However, curren…
Every Step Counts: Step-Level Credit Assignment for Tool-Integrated Text-to-SQL
Yaxun Dai, Baolin Sun, Junying Wang +6
Tool-integrated Text-to-SQL parsing has emerged as a promising paradigm, framing SQL generation as a sequential decision-making process interleaved with tool execution. However, ex…
VisionDirector: Vision-Language Guided Closed-Loop Refinement for Generative Image Synthesis
Meng Chu, Senqiao Yang, Haoxuan Che +8
Generative models can now produce photorealistic imagery, yet they still struggle with the long, multi-goal prompts that professional designers issue. To expose this gap and better…
SearchGym: Bootstrapping Real-World Search Agents via Cost-Effective and High-Fidelity Environment Simulation
Xichen Zhang, Ziyi He, Yinghao Zhu +6
Search agents have emerged as a pivotal paradigm for solving open-ended, knowledge-intensive reasoning tasks. However, training these agents via Reinforcement Learning (RL) faces a…
TraveLLaMA: A Multimodal Travel Assistant with Large-Scale Dataset and Structured Reasoning
Meng Chu, Yukang Chen, Haokun Gui +3
Tourism and travel planning increasingly rely on digital assistance, yet existing multimodal AI systems often lack specialized knowledge and contextual understanding of urban envir…
Agentar-Scale-SQL: Advancing Text-to-SQL through Orchestrated Test-Time Scaling
Pengfei Wang, Baolin Sun, Xuemei Dong +7
State-of-the-art (SOTA) Text-to-SQL methods still lag significantly behind human experts on challenging benchmarks like BIRD. Current approaches that explore test-time scaling lack…