5 papers
FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries
Qijie You, Hao Liang, Mingrui Chen +4
As video becomes increasingly central to information dissemination and multimodal large language models (MLLMs) continue to advance, evaluating video retrieval has become increasin…
TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos
Hengyi Feng, Hao Liang, Mingrui Chen +6
Real-world audio-visual understanding requires chaining evidence that is sparse, temporally dispersed, and split across the visual and auditory streams, whereas existing benchmarks…
DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models
Hao Liang, Zhengyang Zhao, Meiyi Qiang +22
Data-centric training has emerged as a promising direction for improving large language models (LLMs) by optimizing not only model parameters but also the selection, composition, a…
SciAgent: A Unified Multi-Agent System for Generalistic Scientific Reasoning
Xuchen Li, Ruitao Wu, Xuanbo Liu +17
Recent advances in large language models have enabled AI systems to achieve expert-level performance on domain-specific scientific tasks, yet these systems remain narrow and handcr…
MorphoBench: A Benchmark with Difficulty Adaptive to Model Reasoning
Xukai Wang, Xuanbo Liu, Mingrui Chen +16
With the advancement of powerful large-scale reasoning models, effectively evaluating the reasoning capabilities of these models has become increasingly important. However, existin…