4 papers
OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets
Rheeya Uppaal, Seungwoo Lyu, Selina Sung +1
Safe completion requires models to provide useful assistance without enabling harm, but this behavior is difficult to evaluate with isolated prompts. We introduce OpenSafeIntent, a…
HarmThoughts: A Benchmark for Fine-Grained Harmful Behavior Detection in Reasoning Traces
Ishita Kakkar, Enze Zhang, Rheeya Uppaal +1
Large reasoning models (LRMs) produce complex, multi-step reasoning traces, yet safety evaluation remains focused on final outputs, overlooking how harm emerges during reasoning. W…
Shattered Compositionality: Counterintuitive Learning Dynamics of Transformers for Arithmetic
Xingyu Zhao, Darsh Sharma, Rheeya Uppaal +1
Large language models (LLMs) often achieve strong benchmark accuracy yet remain brittle under small distribution shifts. While recent mechanistic studies reveal the discrepancy bet…
Journey Before Destination: On the importance of Visual Faithfulness in Slow Thinking
Rheeya Uppaal, Phu Mon Htut, Min Bai +3
Reasoning-augmented vision language models (VLMs) generate explicit chains of thought that promise greater capability and transparency but also introduce new failure modes: models…