10 papers
What the Reranker Sees: Multi-Aspect Page Annotation for Long-Document Multimodal Question Answering
Guanchen Wu, Jiayuan Ding, Subhabrata Mukherjee +1
Long-document visual question answering (VQA) over documents of tens to hundreds of pages mixing text, tables, charts, and figures typically follows retrieve-then-read pipelines. I…
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…
Crafting Reversible SFT Behaviors in Large Language Models
Yuping Lin, Pengfei He, Yue Xing +5
Supervised fine-tuning (SFT) induces new behaviors in large language models, yet imposes no structural constraint on how these behaviors are distributed within the model. Existing…
HEART: A Unified Benchmark for Assessing Humans and LLMs in Emotional Support Dialogue
Laya Iyer, Kriti Aggarwal, Sanmi Koyejo +3
Supportive conversation depends on skills that go beyond language fluency, including reading emotions, adjusting tone, and navigating moments of resistance, frustration, or distres…
Perfecting Human-AI Interaction at Clinical Scale. Turning Production Signals into Safer, More Human Conversations
Subhabrata Mukherjee, Markel Sanz Ausin, Kriti Aggarwal +24
Healthcare conversational AI agents shouldn't be optimized only for clean benchmark accuracy in production-first regime; they must be optimized for the lived reality of patient con…
GraphGhost: Tracing Structures Behind Large Language Models
Xinnan Dai, Xianxuan Long, Chung-Hsiang Lo +4
Large Language Models (LLMs) exhibit strong reasoning capabilities on structured tasks, yet the internal mechanisms underlying such behaviors remain poorly understood. Existing int…