10 papers
Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models
Aaron Branson Cigres Li, Zhaowei Wang, Yu Zhao +9
Large vision-language models increasingly rely on long-context modeling to reason over documents, hour-level videos, and long-horizon agent trajectories, requiring them to locate r…
SciLens: Multi-modal Scientific Claim Verification with Agentic Entailment and Grounding
Yueming Wang, Tianshi Zheng, Jiaxin Bai +3
Scientific discovery increasingly relies on automated systems that generate hypotheses, inspect multimodal evidence, and validate claims at scale. Yet scientific claim verification…
is Theoretically Large Enough for Embedding-based Top- Retrieval
Zihao Wang, Hang Yin, Lihui Liu +4
This paper studies the Minimal Embeddable Dimension (MED): the least dimension in which there exists a configuration of object vectors so that every subset of size at most …
MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models
Xiyu Ren, Zhaowei Wang, Yiming Du +11
Memory is essential for large vision-language models (LVLMs) to handle long, multimodal interactions, with two method directions providing this capability: long-context LVLMs and m…
MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence
Woohyeon Park, Jaeik Kim, Sunghwan Steve Cho +7
Lesion detection, symptom tracking, and visual explainability are central to real-world medical image analysis, yet current medical Vision-Language Models (VLMs) still lack mechani…
NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM Agents
Tianshi Zheng, Kelvin Kiu-Wai Tam, Newt Hue-Nam K. Nguyen +10
Large language models are emerging as powerful tools for scientific law discovery, a foundational challenge in AI-driven science. However, existing benchmarks for this task suffer…