activity
20242026
most citedFIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights

1 citations · 1 across the 10 of their papers we have counts for

collaborators
Showing cs.CLShow all

21 papers · 1 filter

cs.CL2026

Probing Multimodal Large Language Models on Cognitive Biases in Chinese Short-Video Misinformation

Jen-tse Huang, Chang Chen, Shiyang Lai +3

Short-video platforms have become major channels for misinformation, where deceptive claims frequently leverage visual experiments and social cues. While Multimodal Large Language…

cs.CL2026

Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs

Kaiser Sun, Xiaochuang Yuan, Hongjun Liu +4

Multimodal large language models (MLLMs) can process text presented as images, yet they often perform worse than when the same content is provided as textual tokens. We systematica…

cs.CL2026

Artificial Intolerance: Stigmatizing Language in Clinical Documentation Skews Large Language Model Decision-Making

Jen-tse Huang, Didi Zhou, Faith Kamau +5

Large Language Models (LLMs) are increasingly deployed in high-stakes domains such as clinical decision support and medical documentation. However, the robustness of these models a…

cs.CL2026

Weird Generalization is Weirdly Brittle

Miriam Wanner, Hannah Collison, William Jurayj +3

Weird generalization is a phenomenon in which models fine-tuned on data from a narrow domain (e.g. insecure code) develop surprising traits that manifest even outside that domain (…

cs.CL2026

Task Matters: Knowledge Requirements Shape LLM Responses to Context-Memory Conflict

Kaiser Sun, Fan Bai, Mark Dredze

Large language models (LLMs) draw on both contextual information and parametric memory, yet these sources can conflict. Prior studies have largely examined this issue in contextual…

cs.CL20268 cited

Benchmarking Large Language Models on Answering and Explaining Challenging Medical Questions

Hanjie Chen, Zhouxiang Fang, Yash Singla +1

LLMs have demonstrated impressive performance in answering medical questions, such as achieving passing scores on medical licensing examinations. However, medical board exams or ge…