2 papers
cs.CV2026
Look Where It Counts: A Free, Label-Free Visual Evidence Signal for Fine-Grained Vision-Language Reasoning
Santi Ram Tiwari, Nihal Naik, Devbrat Pandey +1
Multimodal large language models (MLLMs) fail at fine-grained visual questions less because they cannot reason than because they never see the evidence: high-resolution images are…
cs.CV2026
ChronoStitch: Training-Free Composition of Visual KV Memories for Long-Horizon Temporal Reasoning
Santiram Tiwari, Nishant Sinha, Kunal Kislay
Long-video question answering requires a model to preserve visual evidence over time without repeatedly reprocessing the same video. A practical approach is to store the vision-lan…