18 papers
Structurally-bounded Agentic Graph Exploration for Evidence-Grounded Scholarly DeepSearch
Rima Hazra, Sayan Layek, Somnath Banerjee +2
We present Crase, a bounded and inspectable alternative to deep research agents for scholarly search. Instead of an open-ended search loop, Crase queries a search engine once for s…
Lost in Interpretation: The Plausibility-Faithfulness Trade-off in Cross-Lingual Explanations
Somnath Banerjee, Pranav Jha, Rima Hazra +1
LLMs deployed multilingually are often audited via English explanations for non-English inputs. We evaluate extractive explanations ''where the model identifies input token spans a…
SafeMath: Inference-time Safety improves Math Accuracy
Sagnik Basu, Subhrajit Mitra, Aman Juneja +3
Recent research points toward LLMs being manipulated through adversarial and seemingly benign inputs, resulting in harmful, biased, or policy-violating outputs. In this paper, we s…
SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems
Rima Hazra, Bikram Ghuku, Ilona Marchenko +5
Large language models are rapidly being deployed as AI tutors, yet current evaluation paradigms assess problem-solving accuracy and generic safety in isolation, failing to capture…
Bridging the Multilingual Safety Divide: Efficient, Culturally-Aware Alignment for Global South Languages
Somnath Banerjee, Rima Hazra, Animesh Mukherjee
Large language models (LLMs) are being deployed across the Global South, where everyday use involves low-resource languages, code-mixing, and culturally specific norms. Yet safety…
From Fluent to Verifiable: Claim-Level Auditability for Deep Research Agents
Razeen A Rasheed, Somnath Banerjee, Animesh Mukherjee +1
A deep research agent produces a fluent scientific report in minutes; a careful reader then tries to verify the main claims and discovers the real cost is not reading, but tracing:…