collaborators

5 papers

cs.CL2026

Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+

Sherzod Hakimov, Karl Osswald, Jelle Psurek +3

We evaluate large language models (LLMs) as language agents playing goal-directed dialogue games in self-play across 30 languages: the 24 official EU languages plus six others. Unl…

cs.CL2026

TrustMargin: Training-Free Arbitration between Parametric Memory and Retrieved Evidence in Large Language Models

Jingyan Xu, Hong Shi, Yi Shan +4

Large language models answer knowledge-intensive questions using both parametric memory and retrieved evidence, but neither source is uniformly reliable. Retrieval can fill knowled…

cs.AI2026

SGR-Bench: Benchmarking Search Agents on State-Gated Retrieval

Ningyuan Li, Haiyang Shen, Mugeng Liu +4

Recent advances in large language models and tool-using agents have expanded the range of benchmarked web tasks. Yet an important class of specialized retrieval tasks remains under…

cs.CL2026

CoAuthorAI: A Human in the Loop System For Scientific Book Writing

Yangjie Tian, Xungang Gu, Yun Zhao +8

Large language models (LLMs) are increasingly used in scientific writing but struggle with book-length tasks, often producing inconsistent structure and unreliable citations. We in…

cs.CL2024

Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Yun He, Di Jin, Chaoqi Wang +16

Large Language Models (LLMs) have demonstrated impressive capabilities in various tasks, including instruction following, which is crucial for aligning model outputs with user expe…