Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Consistency of Large Reasoning Models Under Multi-Turn Attacks
Yubo Li, Ramayya Krishnan, Rema Padman
Large reasoning models with reasoning capabilities achieve state-of-the-art performance on complex tasks, but their robustness under multi-turn adversarial pressure remains underex…
cs.AI2025
Co-Alignment: Rethinking Alignment as Bidirectional Human-AI Cognitive Adaptation
Yubo Li, Weiyi Song
Current AI alignment through RLHF follows a single directional paradigm that AI conforms to human preferences while treating human cognition as fixed. We propose a shift to co-alig…