3 papers
cs.CL2025
EvalCards: A Framework for Standardized Evaluation Reporting
Ruchira Dhar, Danae Sanchez Villegas, Antonia Karamolegkou +11
Evaluation has long been a central concern in NLP, and transparent reporting practices are more critical than ever in today's landscape of rapidly released open-access models. Draw…
cs.AI2025
What if Othello-Playing Language Models Could See?
Xinyi Chen, Yifei Yuan, Jiaang Li +3
Language models are often said to face a symbol grounding problem. While some have argued the problem can be solved without resort to other modalities, many have speculated that gr…
cs.CL2025
Lost at the Beginning of Reasoning
Baohao Liao, Xinyi Chen, Sara Rajaee +5
Recent advancements in large language models (LLMs) have significantly advanced complex reasoning capabilities, particularly through extended chain-of-thought (CoT) reasoning that…