activity
20242026
most citedControl Illusion: The Failure of Instruction Hierarchies in Large Language Models

2 citations · 2 across the 10 of their papers we have counts for

collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL2026

SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs

Debopriyo Banerjee, Kapil Rajesh Kavitha, Angana Borah +11

Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguistic diversity and culturally…

cs.CL2026

Jais 2: A Family of Arabic-Centric Open Large Language Models

Mohamed Anwar, Abed Alhakim Freihat, George Ibrahim +57

Jais 2 is a family of Arabic-centric large language models developed jointly by MBZUAI, Cerebras, and Inception, designed to advance Arabic-centric language modeling, with strong p…

cs.CL2026

From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves

Haritz Puerto, Haonan Li, Xudong Han +2

Large reasoning models (LRMs) produce reasoning traces (RTs) that often contain sensitive information. These leaky thoughts are difficult to control and frequently violate explicit…

cs.CL2025

Llama-3-Nanda-10B-Chat: An Open Generative Large Language Model for Hindi

Monojit Choudhury, Shivam Chauhan, Rocktim Jyoti Das +27

Developing high-quality large language models (LLMs) for moderately resourced languages presents unique challenges in data availability, model adaptation, and evaluation. We introd…

cs.CL20252 cited

Control Illusion: The Failure of Instruction Hierarchies in Large Language Models

Yilin Geng, Haonan Li, Honglin Mu +5

Large language models (LLMs) are increasingly deployed with hierarchical instruction schemes, where certain instructions (e.g., system-level directives) are expected to take preced…

cs.CL2025

SCALAR: Scientific Citation-based Live Assessment of Long-context Academic Reasoning

Renxi Wang, Honglin Mu, Liqun Ma +5

Long-context understanding has emerged as a critical capability for large language models (LLMs). However, evaluating this ability remains challenging. We present SCALAR, a benchma…