activity
20222026
collaborators

11 papers

cs.CL2026

Last Translation Benchmark

Vilém Zouhar, Niyati Bafna, Mukund Choudhary +241

For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, stan…

cs.AI2026

When Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs

Zhengxiang Wang, Owen Rambow

Visual grounding is typically evaluated as a one-shot mapping from an informative referring expression to a visual target. This formulation misses a central property of real-world…

cs.CL2026

LVLMs and Humans Ground Differently in Referential Communication

Peter Zeng, Weiling Li, Amie J. Paige +6

For generative AI agents to partner effectively with human users, the ability to accurately predict human intent is critical. But this ability to collaborate remains limited by a c…

cs.CL2026

Measuring Iterative Temporal Reasoning with Time Puzzles

Zhengxiang Wang, Zeyu Dong

Tool use, such as web search, has become a standard capability even in freely available large language models (LLMs). However, existing benchmarks evaluate temporal reasoning mainl…

cs.CL2025

LVLMs are Bad at Overhearing Human Referential Communication

Zhengxiang Wang, Weiling Li, Panagiotis Kaliosis +2

During spontaneous conversations, speakers collaborate on novel referring expressions, which they can then re-use in subsequent conversations. Understanding such referring expressi…

cs.CL2025

Catch Me If You Can? Not Yet: LLMs Still Struggle to Imitate the Implicit Writing Styles of Everyday Authors

Zhengxiang Wang, Nafis Irtiza Tripto, Solha Park +2

As large language models (LLMs) become increasingly integrated into personal writing tools, a critical question arises: can LLMs faithfully imitate an individual's writing style fr…