6 papers
QEvict: Recoverable Quantized KV Eviction for Attention-Drift-Robust Long-Context Decoding
Ayushman Garg, Akshita Gupta, Shaswata Bhattacharya +3
Autoregressive large language model inference is increasingly constrained by the memory footprint of the Key-Value (KV) cache. A dominant line of work reduces this footprint by evi…
ReCap: Lightweight Referential Grounding for Coherent Story Visualization
Aditya Arora, Akshita Gupta, Pau Rodriguez +1
Story Visualization aims to generate a sequence of images that faithfully depicts a textual narrative that preserve character identity, spatial configuration, and stylistic coheren…
HaloProbe: Bayesian Detection and Mitigation of Object Hallucinations in Vision-Language Models
Reihaneh Zohrabi, Hosein Hasani, Akshita Gupta +3
Large vision-language models can produce object hallucinations in image descriptions, highlighting the need for effective detection and mitigation strategies. Prior work commonly r…
A multi-modal dataset for insect biodiversity with imagery and DNA at the trap and individual level
Johanna Orsholm, John Quinto, Hannu Autto +26
Insects comprise millions of species, many experiencing severe population declines under environmental and habitat changes. High-throughput approaches are crucial for accelerating…
LoSA: Long-Short-range Adapter for Scaling End-to-End Temporal Action Localization
Akshita Gupta, Gaurav Mittal, Ahmed Magooda +3
Temporal Action Localization (TAL) involves localizing and classifying action snippets in an untrimmed video. The emergence of large video foundation models has led RGB-only video…
Open-Vocabulary Temporal Action Localization using Multimodal Guidance
Akshita Gupta, Aditya Arora, Sanath Narayan +3
Open-Vocabulary Temporal Action Localization (OVTAL) enables a model to recognize any desired action category in videos without the need to explicitly curate training data for all…