11 papers
Death of the Novel(ty): Beyond n-Gram Novelty as a Metric for Textual Creativity
Arkadiy Saakyan, Najoung Kim, Smaranda Muresan +1
N-gram novelty is widely used to evaluate language models' ability to generate text outside of their training data. More recently, it has also been adopted as a metric for measurin…
Designing and Evaluating Chain-of-Hints for Scientific Question Answering
Anubhav Jangra, Smaranda Muresan
LLMs are reshaping education, with students increasingly relying on them for learning. Implemented using general-purpose models, these systems are likely to give away the answers,…
LiveNewsBench: Evaluating LLM Web Search Capabilities with Freshly Curated News
Yunfan Zhang, Kathleen McKeown, Smaranda Muresan
Large Language Models (LLMs) with agentic web search capabilities show strong potential for tasks requiring real-time information access and complex fact retrieval, yet evaluating…
LLMs as Science Journalists: Supporting Early-stage Researchers in Communicating Their Science to the Public
Milad Alshomary, Grace Li, Anubhav Jangra +3
The scientific community needs tools that help early-stage researchers effectively communicate their findings and innovations to the public. Although existing general-purpose Large…
XAM: Interactive Explainability for Authorship Attribution Models
Milad Alshomary, Anisha Bhatnagar, Peter Zeng +3
We present IXAM, an Interactive eXplainability framework for Authorship Attribution Models. Given an authorship attribution (AA) task and an embedding-based AA model, our tool enab…
VideoNorms: Benchmarking Cultural Awareness of Video Language Models
Nikhil Reddy Varimalla, Yunfei Xu, Arkadiy Saakyan +2
As Video Large Language Models (VideoLLMs) are deployed globally, it is important to assess their ability to reason across cultural contexts. To advance cultural norm awareness eva…