1 citations · 1 across the 1 of their papers we have counts for
4 papers
When Benchmarks Age: Temporal Misalignment through Large Language Model Factuality Evaluation
Xunyi Jiang, Dingyi Chang, Julian McAuley +1
The rapid evolution of large language models (LLMs) and the real world has outpaced the static nature of widely used evaluation benchmarks, raising concerns about their reliability…
WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music Reasoning
Gagan Mundada, Yash Vishe, Amit Namburi +4
Recent advances in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities across various vision-language tasks. However, their reasoning abilities in th…
Improving In-Context Learning with Reasoning Distillation
Nafis Sadeq, Xin Xu, Zhouhang Xie +4
Language models rely on semantic priors to perform in-context learning, which leads to poor performance on tasks involving inductive reasoning. Instruction-tuning methods based on…
BiasEdit: Debiasing Stereotyped Language Models via Model Editing
Xin Xu, Wei Xu, Ningyu Zhang +1
Previous studies have established that language models manifest stereotyped biases. Existing debiasing strategies, such as retraining a model with counterfactual data, representati…