activity
20242026
collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2026

SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs

Debopriyo Banerjee, Kapil Rajesh Kavitha, Angana Borah +11

Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguistic diversity and culturally…

cs.CL2026

From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves

Haritz Puerto, Haonan Li, Xudong Han +2

Large reasoning models (LRMs) produce reasoning traces (RTs) that often contain sensitive information. These leaky thoughts are difficult to control and frequently violate explicit…

cs.CL2026

SCALAR: Scientific Citation-based Live Assessment of Long-context Academic Reasoning

Renxi Wang, Honglin Mu, Liqun Ma +5

Long-context understanding has emerged as a critical capability for large language models (LLMs). However, evaluating this ability remains challenging. We present SCALAR, a benchma…

cs.CL20252 cited

Control Illusion: The Failure of Instruction Hierarchies in Large Language Models

Yilin Geng, Haonan Li, Honglin Mu +5

Large language models (LLMs) are increasingly deployed with hierarchical instruction schemes, where certain instructions (e.g., system-level directives) are expected to take preced…

cs.CL2025

Llama-3-Nanda-10B-Chat: An Open Generative Large Language Model for Hindi

Monojit Choudhury, Shivam Chauhan, Rocktim Jyoti Das +27

Developing high-quality large language models (LLMs) for moderately resourced languages presents unique challenges in data availability, model adaptation, and evaluation. We introd…

cs.CL2025

Stealthy Jailbreak Attacks on Large Language Models via Benign Data Mirroring

Honglin Mu, Han He, Yuxin Zhou +9

Large language model (LLM) safety is a critical issue, with numerous studies employing red team testing to enhance model security. Among these, jailbreak methods explore potential…