7 papers
CDR-Bench: Evaluating Faithful Execution of Compositional, Order-Sensitive Data Refinement Recipes
Yuchen Huang, Xiang Li, Zhenqing Ling +5
Data refinement involves executing multi-step recipes over evolving text states, where both composition and execution order of processing operators determine the outcome. While exi…
Towards Self-Evolving Benchmarks: Synthesizing Agent Trajectories via Test-Time Exploration under Validate-by-Reproduce Paradigm
Dadi Guo, Tianyi Zhou, Dongrui Liu +8
Recent advances in large language models (LLMs) and agent system designs have empowered agents with unprecedented levels of capability. However, existing agent benchmarks are showi…
Diversity-Enhanced Reasoning for Subjective Questions
Yumeng Wang, Zhiyuan Fan, Jiayu Liu +2
Large Reasoning Models (LRMs) with long chain-of-thought capabilities, optimized via reinforcement learning with verifiable rewards (RLVR), excel at objective reasoning tasks like…
Empowering Reliable Visual-Centric Instruction Following in MLLMs
Weilei He, Feng Ju, Zhiyuan Fan +3
Evaluating the instruction-following (IF) capabilities of Multimodal Large Language Models (MLLMs) is essential for rigorously assessing how faithfully model outputs adhere to user…
Environment Scaling for Interactive Agentic Experience Collection: A Survey
Yuchen Huang, Sijia Li, Minghao Liu +5
LLM-based agents can autonomously accomplish complex tasks across various domains. However, to further cultivate capabilities such as adaptive behavior and long-term decision-makin…
CantoASR: Prosody-Aware ASR-LALM Collaboration for Low-Resource Cantonese
Dazhong Chen, Yi-Cheng Lin, Yuchen Huang +5
Automatic speech recognition (ASR) is critical for language accessibility, yet low-resource Cantonese remains challenging due to limited annotated data, six lexical tones, tone san…