27 citations · 46 across the 10 of their papers we have counts for
8 papers
What the Reranker Sees: Multi-Aspect Page Annotation for Long-Document Multimodal Question Answering
Guanchen Wu, Jiayuan Ding, Subhabrata Mukherjee +1
Long-document visual question answering (VQA) over documents of tens to hundreds of pages mixing text, tables, charts, and figures typically follows retrieve-then-read pipelines. I…
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
Xiaomin Li, Yuexing Hao, Jianheng Hou +90
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and inter…
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…
Perfecting Human-AI Interaction at Clinical Scale. Turning Production Signals into Safer, More Human Conversations
Subhabrata Mukherjee, Markel Sanz Ausin, Kriti Aggarwal +24
Healthcare conversational AI agents shouldn't be optimized only for clean benchmark accuracy in production-first regime; they must be optimized for the lived reality of patient con…
TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use
Pengfei He, Zhenwei Dai, Bing He +10
Large language model (LLM)-based agents increasingly rely on tool use to complete real-world tasks. While existing works evaluate the LLMs' tool use capability, they largely focus…
GraphGhost: Tracing Structures Behind Large Language Models
Xinnan Dai, Xianxuan Long, Chung-Hsiang Lo +4
Large Language Models (LLMs) exhibit strong reasoning capabilities on structured tasks, yet the internal mechanisms underlying such behaviors remain poorly understood. Existing int…