collaborators

5 papers

cs.CV2026

MVEB: Massive Video Embedding Benchmark

Adnan El Assadi, Roman Solomatin, Isaac Chung +13

We introduce the Massive Video Embedding Benchmark (MVEB), a 23-task benchmark for video embeddings spanning classification, zero-shot classification, clustering, pair classificati…

cs.CL2026

The Silent Vote: Improving Zero-Shot LLM Reliability by Aggregating Semantic Neighborhoods

Sanket Badhe, Priyanka Tiwari, Deep Shah

Large Language Models are increasingly used as zero-shot classifiers in complex reasoning tasks. However, standard constrained decoding suffers from a phenomenon we define as Renor…

cs.CL2026

CROP: Token-Efficient Reasoning in Large Language Models via Regularized Prompt Optimization

Deep Shah, Sanket Badhe, Nehal Kathrotia +1

Large Language Models utilizing reasoning techniques improve task performance but incur significant latency and token costs due to verbose generation. Existing automatic prompt opt…

cs.CL2026

Long-Tail Knowledge in Large Language Models: Taxonomy, Mechanisms, Interventions and Implications

Sanket Badhe, Deep Shah, Nehal Kathrotia

Large language models (LLMs) are trained on web-scale corpora that exhibit steep power-law distributions, in which the distribution of knowledge is highly long-tailed, with most ap…

cs.IR2026

Taxonomy of the Retrieval System Framework: Pitfalls and Paradigms

Deep Shah, Sanket Badhe, Nehal Kathrotia

Designing an embedding retrieval system requires navigating a complex design space of conflicting trade-offs between efficiency and effectiveness. This work structures these decisi…