collaborators

5 papers

cs.CL2025

EvalCards: A Framework for Standardized Evaluation Reporting

Ruchira Dhar, Danae Sanchez Villegas, Antonia Karamolegkou +11

Evaluation has long been a central concern in NLP, and transparent reporting practices are more critical than ever in today's landscape of rapidly released open-access models. Draw…

cs.AI2025

What if Othello-Playing Language Models Could See?

Xinyi Chen, Yifei Yuan, Jiaang Li +3

Language models are often said to face a symbol grounding problem. While some have argued the problem can be solved without resort to other modalities, many have speculated that gr…

cs.CV2025

RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding

Jiaang Li, Yifei Yuan, Wenyan Li +8

As vision-language models (VLMs) become increasingly integrated into daily life, the need for accurate visual culture understanding is becoming critical. Yet, these models frequent…

cs.CL2025

Revisiting the Othello World Model Hypothesis

Yifei Yuan, Anders Søgaard

Li et al. (2023) used the Othello board game as a test case for the ability of GPT-2 to induce world models, and were followed up by Nanda et al. (2023b). We briefly discuss the or…

cs.IR2024

A Comparative Analysis of Faithfulness Metrics and Humans in Citation Evaluation

Weijia Zhang, Mohammad Aliannejadi, Jiahuan Pei +3

Large language models (LLMs) often generate content with unsupported or unverifiable content, known as "hallucinations." To address this, retrieval-augmented LLMs are employed to i…