activity
20242026
collaborators

8 papers

cs.AI2026

Hunt Instead of Wait: Evaluating Deep Data Research on Large Language Models

Wei Liu, Peijie Yu, Michele Orini +2

The agency expected of Agentic Large Language Models goes beyond answering correctly, requiring autonomy to set goals and decide what to explore. We term this investigatory intelli…

cs.CV2026

Seeing Is Believing? A Benchmark for Multimodal Large Language Models on Visual Illusions and Anomalies

Wenjin Hou, Wei Liu, Han Hu +3

Multimodal Large Language Models (MLLMs) have shown remarkable proficiency on general-purpose vision-language benchmarks, reaching or even exceeding human-level performance. Howeve…

cs.AI2026

Are We Evaluating the Edit Locality of LLM Model Editing Properly?

Wei Liu, Haomei Xu, Hongkai Liu +5

Model editing has recently emerged as a popular paradigm for efficiently updating knowledge in LLMs. A central desideratum of updating knowledge is to balance editing efficacy, i.e…

cs.CL2026

RM-Distiller: Exploiting Generative LLM for Reward Model Distillation

Hongli Zhou, Hui Huang, Wei Liu +8

Reward models (RMs) play a pivotal role in aligning large language models (LLMs) with human preferences. Due to the difficulty of obtaining high-quality human preference annotation…

cs.CR2025

SecReEvalBench: A Multi-turned Security Resilience Evaluation Benchmark for Large Language Models

Huining Cui, Wei Liu

The increasing deployment of large language models in security-sensitive domains necessitates rigorous evaluation of their resilience against adversarial prompt-based attacks. Whil…

cs.CL2025

UQ: Assessing Language Models on Unsolved Questions

Fan Nie, Ken Ziyu Liu, Zihao Wang +11

Benchmarks shape progress in AI research. A useful benchmark should be both difficult and realistic: questions should challenge frontier models while also reflecting real-world usa…