most citedUQ: Assessing Language Models on Unsolved Questions

1 citations · 1 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CV2026

Seeing Is Believing? A Benchmark for Multimodal Large Language Models on Visual Illusions and Anomalies

Wenjin Hou, Wei Liu, Han Hu +3

Multimodal Large Language Models (MLLMs) have shown remarkable proficiency on general-purpose vision-language benchmarks, reaching or even exceeding human-level performance. Howeve…

cs.AI2026

Are We Evaluating the Edit Locality of LLM Model Editing Properly?

Wei Liu, Haomei Xu, Hongkai Liu +5

Model editing has recently emerged as a popular paradigm for efficiently updating knowledge in LLMs. A central desideratum of updating knowledge is to balance editing efficacy, i.e…

cs.CL20251 cited

UQ: Assessing Language Models on Unsolved Questions

Fan Nie, Ken Ziyu Liu, Zihao Wang +11

Benchmarks shape progress in AI research. A useful benchmark should be both difficult and realistic: questions should challenge frontier models while also reflecting real-world usa…

cs.CL2025

XRAG: Cross-lingual Retrieval-Augmented Generation

Wei Liu, Sony Trenous, Leonardo F. R. Ribeiro +2

We propose XRAG, a novel benchmark designed to evaluate the generation abilities of LLMs in cross-lingual Retrieval-Augmented Generation (RAG) settings where the user language does…

cs.CR2025

SecReEvalBench: A Multi-turned Security Resilience Evaluation Benchmark for Large Language Models

Huining Cui, Wei Liu

The increasing deployment of large language models in security-sensitive domains necessitates rigorous evaluation of their resilience against adversarial prompt-based attacks. Whil…