From the 1 of 12 linked papers with an AI index.
12 papers
Clarify-Then-Search: A Clarification Benchmark for Deep Search with End-to-End Nugget Restoration
Deqiang Huang, Jingbo Zhou, Xinjiang Lu +3
Deep search is brittle on underspecified user queries: missing constraints such as time, location, scope, or definitions can lead to retrieval drift and incomplete answers. We intr…
OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
Jingbo Zhou, Yusai Zhao, Qi Bao +12
The paper presents OmegaUse-OfficeVal, a benchmark that evaluates large language model agents on long‑horizon office‑suite tasks while providing economic signals (human labor time…
ATTNPO: Attention-Guided Process Supervision for Efficient Reasoning
Shuaiyi Nie, Siyu Ding, Wenyuan Zhang +7
Large reasoning models trained with reinforcement learning and verifiable rewards (RLVR) achieve strong performance on complex reasoning tasks, yet often overthink, generating redu…
Eureka-Audio: Triggering Audio Intelligence in Compact Language Models
Dan Zhang, Yishu Lei, Jing Hu +10
We present Eureka-Audio, a compact yet high-performance audio language model that achieves competitive performance against models that are 4 to 18 times larger across a broad range…
ERNIE 5.0 Technical Report
Haifeng Wang, Hua Wu, Tian Wu +432
In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…
Extending RLVR to Open-Ended Tasks via Verifiable Multiple-Choice Reformulation
Mengyu Zhang, Siyu Ding, Weichong Yin +2
Reinforcement Learning with Verifiable Rewards(RLVR) has demonstrated great potential in enhancing the reasoning capabilities of large language models (LLMs). However, its success…