10 papers
Beyond Rubrics: Exploration-Guided Evaluation Skills for Reward Modeling
Xing Yue, Linjuan Wu, Daoxin Zhang +2
Open-ended reward modeling requires judges that can follow subtle, domain-specific preferences when verifiable answers are unavailable. Existing rubric-based methods often address…
When Languages Disagree: Self-Evolving Multilingual LLM Judges
Xiyan Fu, Wei Lu
Multilingual LLM-as-a-judge is widely used to evaluate model outputs across languages, but suffers from cross-lingual inconsistency (Fu and Liu, 2025). Existing methods typically t…
Phun-Bench: Evaluating LLMs on Phonological Understanding in Chinese
Xing Yue, Yongliang Shen, Weiming Lu
Language is a vehicle for thought, intricately tied to sounds, symbols, and meaning. However, most large language model (LLM) research focuses on meaning (semantics) and symbols (s…
Beyond Literal Translation: Evaluating Cultural Effectiveness in Social Media UGC
Linjuan Wu, Ruiqi Zhang, Xinze Lyu +7
Social media platforms enable large-scale cross-lingual communication, but translating user-generated content (UGC) remains challenging due to its informal style, cultural referenc…
Self-Distilled Agentic Reinforcement Learning
Zhengxi Lu, Zhiyuan Yao, Zhuowen Han +8
Reinforcement learning (RL) has emerged as a central paradigm for post-training LLM agents, yet its trajectory-level reward signal provides only coarse supervision for long-horizon…
Milestone-Guided Policy Learning for Long-Horizon Language Agents
Zixuan Wang, Yuchen Yan, Hongxing Li +7
While long-horizon agentic tasks require language agents to perform dozens of sequential decisions, training such agents with reinforcement learning remains challenging. We identif…