8 papers
Mitigating Content Effects on Reasoning in Language Models through Fine-Grained Activation Steering
Marco Valentino, Geonhee Kim, Dhairya Dalal +2
Large language models (LLMs) exhibit reasoning biases, often conflating content plausibility with formal logical validity. This can lead to wrong inferences in critical domains, wh…
KidsArtBench: Multi-Dimensional Children's Art Evaluation with Attribute-Aware MLLMs
Mingrui Ye, Chanjin Zheng, Zengyi Yu +4
Multimodal Large Language Models (MLLMs) show remarkable progress across many visual-language tasks; however, their capacity to evaluate artistic expression remains limited. Aesthe…
Position: On the Methodological Pitfalls of Evaluating Base LLMs for Reasoning
Jason Chan, Zhixue Zhao, Robert Gaizauskas
Existing work investigates the reasoning capabilities of large language models (LLMs) to uncover their limitations, human-like biases and underlying processes. Such studies include…
AI-Generated Content in Cross-Domain Applications: Research Trends, Challenges and Propositions
Jianxin Li, Liang Qu, Taotao Cai +13
Artificial Intelligence Generated Content (AIGC) has rapidly emerged with the capability to generate different forms of content, including text, images, videos, and other modalitie…
RULEBREAKERS: Challenging LLMs at the Crossroads between Formal Logic and Human-like Reasoning
Jason Chan, Robert Gaizauskas, Zhixue Zhao
Formal logic enables computers to reason in natural language by representing sentences in symbolic forms and applying rules to derive conclusions. However, in what our study charac…
How Robust is Model Editing after Fine-Tuning? An Empirical Study on Text-to-Image Diffusion Models
Feng He, Zhenyang Liu, Marco Valentino +1
Model editing offers a low-cost technique to inject or correct a particular behavior in a pre-trained model without extensive retraining, supporting applications such as factual co…