activity
20242026
collaborators
Showing cs.CLShow all

18 papers · 1 filter

cs.CL2026

Moral Safety in LLMs: Exposing Performative Compliance with Puzzled Cues

Mohammadamin Shafiei, Shuyue Stella Li, Yulia Tsvetkov

As large language models take on morally consequential roles in healthcare, legal, and hiring contexts, we need to examine whether their ethical behaviors are genuine or superficia…

cs.CL2026

Multilingual Reasoning Cascades Need More Context

Arnav Mazumder, Dengjia Zhang, Shuyue Stella Li +2

Translation cascades for reasoning translate the query from another language to English, reason in English, and translate the answer back to the original language. This is a compet…

cs.CL2026

EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments

Jundong Xu, Qingchuan Li, Jiaying Wu +11

Large language model (LLM) agents have achieved strong performance on a wide range of benchmarks, yet most evaluations assume static environments. In contrast, real-world deploymen…

cs.CL2026

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments

Zhiyuan Zeng, Hamish Ivison, Yiping Wang +14

We introduce Reinforcement Learning (RL) with Adaptive Verifiable Environments (RLVE), an approach using verifiable environments that procedurally generate problems and provide alg…

cs.CL2026

Deep Reasoning in General Purpose Agents via Structured Meta-Cognition

Dean Light, Michael Theologitis, Kshitish Ghate +7

Humans intuitively solve complex problems by flexibly shifting among reasoning modes: they plan, execute, revise intermediate goals, resolve ambiguity through associative judgment,…

cs.CL2026

HorizonBench: Long-Horizon Personalization with Evolving Preferences

Shuyue Stella Li, Bhargavi Paranjape, Kerem Oktar +9

User preferences evolve across months of interaction, and tracking them requires inferring when a stated preference has been changed by a subsequent life event. We define this prob…