2 papers
cs.CL2026
HarmThoughts: A Benchmark for Fine-Grained Harmful Behavior Detection in Reasoning Traces
Ishita Kakkar, Enze Zhang, Rheeya Uppaal +1
Large reasoning models (LRMs) produce complex, multi-step reasoning traces, yet safety evaluation remains focused on final outputs, overlooking how harm emerges during reasoning. W…
cs.CL2025
Fluent but Unfeeling: The Emotional Blind Spots of Language Models
Bangzhao Shu, Isha Joshi, Melissa Karnaze +5
The versatility of Large Language Models (LLMs) in natural language understanding has made them increasingly popular in mental health research. While many studies explore LLMs' cap…