activity
20242026
collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2026

Beyond Rubrics: Exploration-Guided Evaluation Skills for Reward Modeling

Xing Yue, Linjuan Wu, Daoxin Zhang +2

Open-ended reward modeling requires judges that can follow subtle, domain-specific preferences when verifiable answers are unavailable. Existing rubric-based methods often address…

cs.CL2026

When Languages Disagree: Self-Evolving Multilingual LLM Judges

Xiyan Fu, Wei Lu

Multilingual LLM-as-a-judge is widely used to evaluate model outputs across languages, but suffers from cross-lingual inconsistency (Fu and Liu, 2025). Existing methods typically t…

cs.CL2026

Phun-Bench: Evaluating LLMs on Phonological Understanding in Chinese

Xing Yue, Yongliang Shen, Weiming Lu

Language is a vehicle for thought, intricately tied to sounds, symbols, and meaning. However, most large language model (LLM) research focuses on meaning (semantics) and symbols (s…

cs.CL2026

Beyond Literal Translation: Evaluating Cultural Effectiveness in Social Media UGC

Linjuan Wu, Ruiqi Zhang, Xinze Lyu +7

Social media platforms enable large-scale cross-lingual communication, but translating user-generated content (UGC) remains challenging due to its informal style, cultural referenc…

cs.CL2026

Milestone-Guided Policy Learning for Long-Horizon Language Agents

Zixuan Wang, Yuchen Yan, Hongxing Li +7

While long-horizon agentic tasks require language agents to perform dozens of sequential decisions, training such agents with reinforcement learning remains challenging. We identif…

cs.CL2024

ProSwitch: Knowledge-Guided Instruction Tuning to Switch Between Professional and Non-Professional Responses

Chang Zong, Yuyan Chen, Weiming Lu +4

Large Language Models (LLMs) have demonstrated efficacy in various linguistic applications, including question answering and controlled text generation. However, studies into their…