Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Frames2LoRA: Parametric Video Internalization for Vision-Language Models
Manan Suri, Sarvesh Baskar, Dinesh Manocha
Processing video in vision-language models is expensive: each frame occupies hundreds of tokens, and inference cost scales with every frame and every repeated query. We introduce F…
cs.CV2026
DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams
Anirudh Iyengar Kaniyar Narayana Iyengar, Tampu Ravi Kumar, Gaurav Najpande +4
Diagram question answering (DQA) requires models to interpret structured visual representations such as charts, maps, infographics, circuit schematics, and scientific diagrams. Rec…
cs.CV2026
Learning Illumination Control in Diffusion Models
Nishit Anand, Manan Suri, Christopher Metzler +2
Controlling illumination in images is essential for photography and visual content creation. While closed-source models have demonstrated impressive illumination control, open-sour…