7 papers
MESA:Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory
Beidi Zhao, Yaoqi Chen, Yuru Feng +10
Long-horizon agents accumulate trajectories spanning hundreds of interleaved reasoning, action, and observation steps, where answering a query may depend on evidence buried far bac…
SkillHEX: Improving Agent Skills via Hypothesis-Driven Autonomous Exploration and Exploitation
Yuru Feng, Yaoqi Chen, Beidi Zhao +7
Although agent skills equip LLMs with reusable procedural knowledge, manual maintenance suffers from high costs, unscalability, and misalignment. Real-world deployments thus requir…
When RAG Hurts: Diagnosing and Mitigating Attention Distraction in Retrieval-Augmented LVLMs
Beidi Zhao, Wenlong Deng, Xinting Liao +4
While Retrieval-Augmented Generation (RAG) is one of the dominant paradigms for enhancing Large Vision-Language Models (LVLMs) on knowledge-based VQA tasks, recent work attributes…
UltraVR: A Diagnostic Ultra-Resolution Image-VQA Benchmark for Evidence-Grounded Reasoning
Gexin Huang, Yanting Yang, Myeongkyun Kang +6
Vision-language models (VLMs) excel on visual question answering and multimodal reasoning benchmarks. Yet their capability on ultra-resolution images - where critical evidence is t…
Verification Mirage: Mapping the Reliability Boundary of Self-Verification in Medical VQA
Ruinan Jin, Beidi Zhao, Myeongkyun Kang +2
Self-verification, re-invoking the same vision language model (VLM) in a fresh context to check its own generated answer, is increasingly used as a default safety layer for medical…
Plug-and-Play Logit Fusion for Heterogeneous Pathology Foundation Models
Gexin Huang, Anqi Li, Yusheng Tan +4
Pathology foundation models (FMs) have become central to computational histopathology, offering strong transfer performance across a wide range of diagnostic and prognostic tasks.…