activity
20242026
collaborators

8 papers

cs.AI2026

ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay

Zhexin Hu, Li Wang, Xiaohan Wang +4

Adaptive context compression is vital for scaling Large Language Models (LLMs) to complex, multi-turn agent tasks. However, rule-based compression methods may discard task-critical…

cs.SD2026

VoiceGiraffe: A Benchmark for Extreme Long-Context Audio-Language Understanding

Jashin Ye, Dongxiao Wang, Yixuan Ye +10

While large audio language models (LALMs) have achieved remarkable progress in audio processing at the second- or minute-level scale, understanding hour-level audio remains a funda…

cs.LG2026

When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards

Li Wang, Xiaodong Lu, Xiaohan Wang +5

Large Language Models (LLMs) have achieved remarkable advancements in reasoning capabilities empowered by Reinforcement Learning with Verifiable Rewards (RLVR). Nonetheless, RLVR i…

cs.CL2026

Large Language Models Could Be Rote Learners

Yuyang Xu, Renjun Hu, Haochao Ying +3

Benchmark-based evaluation, e.g., multiple-choice questions (MCQs) and open-ended questions (OEQs), is widely used for evaluating Large Language Models (LLMs), yet their reliabilit…

cs.LG2025

SSPO: Self-traced Step-wise Preference Optimization for Process Supervision and Reasoning Compression

Yuyang Xu, Yi Cheng, Haochao Ying +5

Test-time scaling has proven effective in further enhancing the performance of pretrained Large Language Models (LLMs). However, mainstream post-training methods (i.e., reinforceme…

cs.IR2025

Behavior Modeling Space Reconstruction for E-Commerce Search

Yejing Wang, Chi Zhang, Xiangyu Zhao +8

Delivering superior search services is crucial for enhancing customer experience and driving revenue growth. Conventionally, search systems model user behaviors by combining user p…