12 papers
Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models
Aaron Branson Cigres Li, Zhaowei Wang, Yu Zhao +9
Large vision-language models increasingly rely on long-context modeling to reason over documents, hour-level videos, and long-horizon agent trajectories, requiring them to locate r…
SciLens: Multi-modal Scientific Claim Verification with Agentic Entailment and Grounding
Yueming Wang, Tianshi Zheng, Jiaxin Bai +3
Scientific discovery increasingly relies on automated systems that generate hypotheses, inspect multimodal evidence, and validate claims at scale. Yet scientific claim verification…
Reinforcing Dual-Path Reasoning in Spatial Vision Language Models
Yatai Ji, An-Chieh Cheng, Yang Fu +13
Spatial VLMs have made substantial progress in geometric perception, yet complex spatial reasoning requiring multi-step inference over depth, distance, and scene relations remains…
Context-Driven Incremental Compression for Multi-Turn Dialogue Generation
Yeongseo Jung, Jaehyeok Kim, Eunseo Jung +5
Modern conversational agents condition on an ever-growing dialogue history at each turn, incurring redundant attention and encoding costs that grow with conversation length. Naive…
The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes
Avinash Anand, Mahisha Ramesh, Avni Mittal +8
Reasoning has become central to how Large Language Models (LLMs) are evaluated and interpreted, spanning Chain-of-Thought (CoT), mathematical problem-solving, multi-hop question an…
is Theoretically Large Enough for Embedding-based Top- Retrieval
Zihao Wang, Hang Yin, Lihui Liu +4
This paper studies the Minimal Embeddable Dimension (MED): the least dimension in which there exists a configuration of object vectors so that every subset of size at most …