13 papers
SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Acquisition
Yuqi Tang, Chenyi Zhou, Libin Wang +3
Large language model (LLM) agents have been increasingly adopted in scientific research for organizing and invoking specialized computational tools. However, their reliance on pred…
Beacon: Knowing When and How to Perform Agentic Visual Reasoning
Qixun Wang, Yang Shi, Letian Cheng +11
The paper introduces Beacon, an agentic visual reasoning system that learns when to invoke external tools and how to use them effectively, improving multimodal large language model…
RefCaptioner: Multi-Reference Image-Grounded Video Captioning
Tengfei Liu, Yang Shi, Yuran Wang +16
Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-…
KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation
Yuqi Tang, Tengfei Liu, Yizheng Lai +18
The paper introduces KeyFrame-Compass, a benchmark and evaluation framework for assessing how well video generation models follow supplied keyframes while preserving overall video…
MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
Xiaohan Zhang, Yuqing Wen, Junlin Chen +9
The paper introduces MultiRef-Compass, a benchmark for evaluating models that generate synchronized audio‑video content conditioned on multiple references and textual instructions,…
LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV
Tengfei Liu, Yang Shi, Xuanyu Zhu +17
Audio-visual generation is rapidly advancing from short clips to minute-long content, while existing evaluation protocols remain largely confined to short-form settings. Existing b…