collaborators

6 papers

cs.CR2026

BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure

Shenghan Zheng, Zonglin Di, Yimin Liu +19

LM-agent benchmarks increasingly function as interactive evaluation infrastructure. Agents observe state, call tools, modify workspaces, submit artifacts, and receive rewards from…

cs.CV2026

TDDN: Text-aligned Diffused DINO Network for Puzzle Understanding

Harsha Patnala, Debopriyo Banerjee, Ayush Sunil Munot +1

Structured visual reasoning, such as image puzzles, demands fine-grained visual perception, an ability current Vision Language Models (VLMs) lack. VLMs built on CLIP-based ViT back…

cs.CL2026

Last Translation Benchmark

Vilém Zouhar, Niyati Bafna, Mukund Choudhary +241

For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, stan…

cs.CL2026

SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs

Debopriyo Banerjee, Kapil Rajesh Kavitha, Angana Borah +11

Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguistic diversity and culturally…

cs.CV2026

MVEB: Massive Video Embedding Benchmark

Adnan El Assadi, Roman Solomatin, Isaac Chung +13

We introduce the Massive Video Embedding Benchmark (MVEB), a 23-task benchmark for video embeddings spanning classification, zero-shot classification, clustering, pair classificati…

cs.SD2026

MAEB: Massive Audio Embedding Benchmark

Adnan El Assadi, Isaac Chung, Chenghao Xiao +15

We introduce the Massive Audio Embedding Benchmark (MAEB), a large-scale benchmark covering 30 tasks across speech, music, environmental sounds, and cross-modal audio-text reasonin…