19 papers · 1 filter
Last Translation Benchmark
Vilém Zouhar, Niyati Bafna, Mukund Choudhary +241
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, stan…
Register Shifts Break LLM Safety: A Bengali Benchmark with Culturally Grounded Harms
Naymul Islam, Nusrat Jahan Lia, Shubhashis Roy Dipta +2
Bengali is the seventh-most-spoken language globally, yet LLM safety evaluation remains overwhelmingly English-centric. We introduce BanglaSafe, a benchmark of 879 Bengali prompts…
Cornerstones or Stumbling Blocks? Deciphering the Rock Tokens in On-Policy Distillation
Yuxuan Jiang, Runchao Li, Shubhashis Roy Dipta +2
While recent work in Reinforcement Learning with Verifiable Rewards (RLVR) has shown that a small subset of critical tokens disproportionately drives reasoning gains, an analogous…
DecomposeRL: Learning to Ask Useful, Informative, and Diverse Questions for Semi-Supervised, Traceable Claim Verification
Shubhashis Roy Dipta, Ankur Padia, Francis Ferraro
Claim verification splits between end-to-end classifiers that are accurate but yields no inspectable traces, and decomposition-based methods produce inspectable traces but lag perf…
AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators
Aritra Mazumder, Shubhashis Roy Dipta, Nusrat Jahan Lia +10
Multi-agent systems achieve state-of-the-art outcomes through peer collaboration. However, when an agent in the pipeline silently drops a constraint, the system's final output may…
Many Dialects, Many Languages, One Cultural Lens: Evaluating Multilingual VLMs for Bengali Culture Understanding Across Historically Linked Languages and Regional Dialects
Nurul Labib Sayeedi, Md. Faiyaz Abdullah Sayeedi, Shubhashis Roy Dipta +6
Bangla culture is richly expressed through region, dialect, history, food, politics, media, and everyday visual life, yet it remains underrepresented in multimodal evaluation. To a…