7 papers
ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay
Zhexin Hu, Li Wang, Xiaohan Wang +4
Adaptive context compression is vital for scaling Large Language Models (LLMs) to complex, multi-turn agent tasks. However, rule-based compression methods may discard task-critical…
VoiceGiraffe: A Benchmark for Extreme Long-Context Audio-Language Understanding
Jashin Ye, Dongxiao Wang, Yixuan Ye +10
While large audio language models (LALMs) have achieved remarkable progress in audio processing at the second- or minute-level scale, understanding hour-level audio remains a funda…
When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards
Li Wang, Xiaodong Lu, Xiaohan Wang +5
Large Language Models (LLMs) have achieved remarkable advancements in reasoning capabilities empowered by Reinforcement Learning with Verifiable Rewards (RLVR). Nonetheless, RLVR i…
Large Language Models Could Be Rote Learners
Yuyang Xu, Renjun Hu, Haochao Ying +3
Benchmark-based evaluation, e.g., multiple-choice questions (MCQs) and open-ended questions (OEQs), is widely used for evaluating Large Language Models (LLMs), yet their reliabilit…
SSPO: Self-traced Step-wise Preference Optimization for Process Supervision and Reasoning Compression
Yuyang Xu, Yi Cheng, Haochao Ying +5
Test-time scaling has proven effective in further enhancing the performance of pretrained Large Language Models (LLMs). However, mainstream post-training methods (i.e., reinforceme…
Behavior Modeling Space Reconstruction for E-Commerce Search
Yejing Wang, Chi Zhang, Xiangyu Zhao +8
Delivering superior search services is crucial for enhancing customer experience and driving revenue growth. Conventionally, search systems model user behaviors by combining user p…