activity
20242026
collaborators

7 papers

cs.CV2026

From Visual Cues to Spoken Narration: Rethinking Audio Description

Akshita Gupta, Aditya Arora, Federico Tombari +2

Audio Description (AD) provides spoken narration of visual events during dialogue gaps, making movies accessible to visually impaired audiences. The problem requires determining bo…

cs.LG2026

QEvict: Recoverable Quantized KV Eviction for Attention-Drift-Robust Long-Context Decoding

Ayushman Garg, Akshita Gupta, Shaswata Bhattacharya +3

Autoregressive large language model inference is increasingly constrained by the memory footprint of the Key-Value (KV) cache. A dominant line of work reduces this footprint by evi…

cs.CV2026

ReCap: Lightweight Referential Grounding for Coherent Story Visualization

Aditya Arora, Akshita Gupta, Pau Rodriguez +1

Story Visualization aims to generate a sequence of images that faithfully depicts a textual narrative that preserve character identity, spatial configuration, and stylistic coheren…

cs.CV2026

HaloProbe: Bayesian Detection and Mitigation of Object Hallucinations in Vision-Language Models

Reihaneh Zohrabi, Hosein Hasani, Akshita Gupta +3

Large vision-language models can produce object hallucinations in image descriptions, highlighting the need for effective detection and mitigation strategies. Prior work commonly r…

cs.CV2025

A multi-modal dataset for insect biodiversity with imagery and DNA at the trap and individual level

Johanna Orsholm, John Quinto, Hannu Autto +26

Insects comprise millions of species, many experiencing severe population declines under environmental and habitat changes. High-throughput approaches are crucial for accelerating…

cs.CV2024

Open-Vocabulary Temporal Action Localization using Multimodal Guidance

Akshita Gupta, Aditya Arora, Sanath Narayan +3

Open-Vocabulary Temporal Action Localization (OVTAL) enables a model to recognize any desired action category in videos without the need to explicitly curate training data for all…