activity
20242026
collaborators

23 papers

cs.AI2026

Different Feedback, Different Updates: Selective Self-Learning from User Interactions for Large Language Models

Xuanchen Li, Haitao Li, Yujia Zhou +5

User feedback offers natural supervision for persistent LLM improvement, but a single message may support multiple behavioral changes with different scopes of generalization. We in…

cs.CL2026

Civil Court Simulation with Large Language Models

Yifan Chen, Haitao Li, Kaiyuan Zhang +3

Court simulation bridges legal education and judicial practice, yet human-based simulations are costly and difficult to scale. Large language models (LLMs) offer a scalable alterna…

cs.CL2026

LexRubric: A Rubric-Guided Diagnostic Benchmark for Open-Ended Legal Tasks

Yifan Chen, Haitao Li, Yiran Hu +6

As large language models (LLMs) are increasingly applied to real-world legal tasks, evaluating the reliability of their open-ended legal responses has become essential. These tasks…

eess.AS2026

AnyAudio-Judge: A Dynamic Rubric-Based Benchmark and Evaluator for Audio Instruction Following

Haitao Li, Tian Tan, Yuguang Yang +2

The rapid advancement of instruction-guided audio generation has highlighted the critical need for robust alignment evaluation. Current automated evaluation methods heavily rely on…

cs.CL2026

Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation

Junjie Chen, Yuxi Dong, Haitao Li +6

As large language models (LLMs) are increasingly used for long-form generation, reliably evaluating long-form outputs has become a critical challenge. LLM-as-a-judge offers a scala…

cs.AI2026

MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop

Xuancheng Li, Haitao Li, Yujia Zhou +2

Reinforcement Learning with Verifiable Rewards (RLVR) is widely used to improve reasoning across domains, but outcome-only scalar rewards are often sparse and uninformative. This l…