7 papers
Register Shifts Break LLM Safety: A Bengali Benchmark with Culturally Grounded Harms
Naymul Islam, Nusrat Jahan Lia, Shubhashis Roy Dipta +2
Bengali is the seventh-most-spoken language globally, yet LLM safety evaluation remains overwhelmingly English-centric. We introduce BanglaSafe, a benchmark of 879 Bengali prompts…
AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP
Aritra Mazumder, Nusrat jahan Lia
Tool-using LLM agents are mostly evaluated assuming all tools work. When a tool times out, returns a week-stale value, or has its description poisoned in deployment, the developer…
Learning What Not to Forget: Long-Horizon Agent Memory from a Few Kilobytes of Learning
Nusrat Jahan Lia, Aritra Mazumder
Long-running language-model systems accumulate interaction history that outgrows the context window, so they must continually evict. When an eviction policy drops a load-bearing de…
AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators
Aritra Mazumder, Shubhashis Roy Dipta, Nusrat Jahan Lia +10
Multi-agent systems achieve state-of-the-art outcomes through peer collaboration. However, when an agent in the pipeline silently drops a constraint, the system's final output may…
Cross-Lingual Sentiment Misalignment: Auditing Multilingual Language Models for Inversion Risk, Dialectal Representation, and Affective Stability
Nusrat Jahan Lia, Shubhashis Roy Dipta
Recent advances in multilingual representation learning aim to bridge the performance gap between high- and low-resource languages, yet their ability to preserve affective meaning…
Exploring Cross-Lingual Knowledge Transfer via Transliteration-Based MLM Fine-Tuning for Critically Low-resource Chakma Language
Adity Khisa, Nusrat Jahan Lia, Tasnim Mahfuz Nafis +4
As an Indo-Aryan language with limited available data, Chakma remains largely underrepresented in language models. In this work, we introduce a novel corpus of contextually coheren…