6 papers
ROSE: Benchmarking the Perception-to-Action Gap in Multimodal Models
Yihao Wang, Zijian He, Jie Ren +1
Multimodal large language models (MLLMs) are increasingly expected to act on visual information, yet the same scene may require different actions under different task contexts. How…
How Do Document Parsers Break? Auditing Structural Vulnerability in Document Intelligence
Yue Chen, Yihao Wang, Ziyi Tang +2
Document Layout Analysis (DLA) pipelines provide structured page representations for retrieval-augmented generation, long-document question answering, and other document intelligen…
Bridging the Detection-to-Abstention Gap in Reasoning Models under Insufficient Information
Renjie Gu, Jiaxu Li, Yihao Wang +8
We highlight a failure mode of large reasoning models on questions with insufficient information: models may recognize that a problem is under-specified, yet still continue reasoni…
RotMoLE: Enhancing Mixture of Low-Rank Experts through Rotational Gating Mechanism
Mengyang Sun, Maochuan Dou, Tao Feng +5
While Large Language Models (LLMs) are commonly fine-tuned to handle domain-specific tasks before being applied to vertical applications, adapting them to complex scenarios with di…
ChunQiuTR: Time-Keyed Temporal Retrieval in Classical Chinese Annals
Yihao Wang, Zijian He, Jie Ren +1
Retrieval shapes how language models access and ground knowledge in retrieval-augmented generation (RAG). In historical research, the target is often not an arbitrary relevant pass…
ResAgent: Entropy-based Prior Point Discovery and Visual Reasoning for Referring Expression Segmentation
Yihao Wang, Jusheng Zhang, Ziyi Tang +2
Referring Expression Segmentation (RES) is a core vision-language segmentation task that enables pixel-level understanding of targets via free-form linguistic expressions, supporti…