23 papers
Different Feedback, Different Updates: Selective Self-Learning from User Interactions for Large Language Models
Xuanchen Li, Haitao Li, Yujia Zhou +5
User feedback offers natural supervision for persistent LLM improvement, but a single message may support multiple behavioral changes with different scopes of generalization. We in…
Civil Court Simulation with Large Language Models
Yifan Chen, Haitao Li, Kaiyuan Zhang +3
Court simulation bridges legal education and judicial practice, yet human-based simulations are costly and difficult to scale. Large language models (LLMs) offer a scalable alterna…
LexRubric: A Rubric-Guided Diagnostic Benchmark for Open-Ended Legal Tasks
Yifan Chen, Haitao Li, Yiran Hu +6
As large language models (LLMs) are increasingly applied to real-world legal tasks, evaluating the reliability of their open-ended legal responses has become essential. These tasks…
AnyAudio-Judge: A Dynamic Rubric-Based Benchmark and Evaluator for Audio Instruction Following
Haitao Li, Tian Tan, Yuguang Yang +2
The rapid advancement of instruction-guided audio generation has highlighted the critical need for robust alignment evaluation. Current automated evaluation methods heavily rely on…
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
Junjie Chen, Yuxi Dong, Haitao Li +6
As large language models (LLMs) are increasingly used for long-form generation, reliably evaluating long-form outputs has become a critical challenge. LLM-as-a-judge offers a scala…
MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop
Xuancheng Li, Haitao Li, Yujia Zhou +2
Reinforcement Learning with Verifiable Rewards (RLVR) is widely used to improve reasoning across domains, but outcome-only scalar rewards are often sparse and uninformative. This l…