Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Green Shielding: A User-Centric Approach Towards Trustworthy AI
Aaron J. Li, Nicolas Sanchez, Hao Huang +8
Large language models (LLMs) are increasingly deployed, yet their outputs can be highly sensitive to routine, non-adversarial variation in how users phrase queries, a gap not well…
cs.CL2026
MR-Align: Meta-Reasoning Informed Factuality Alignment for Large Reasoning Models
Xinming Wang, Jian Xu, Bin Yu +9
Large reasoning models (LRMs) show strong capabilities in complex reasoning, yet their marginal gains on evidence-dependent factual questions are limited. We find this limitation i…