Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
ODESteer: A Unified ODE-Based Steering Framework for LLM Alignment
Hongjue Zhao, Haosen Sun, Jiangtao Kong +8
Activation steering, or representation engineering, offers a lightweight approach to align large language models (LLMs) by manipulating their internal activations at inference time…
cs.AI2025
HAICOSYSTEM: An Ecosystem for Sandboxing Safety Risks in Human-AI Interactions
Xuhui Zhou, Hyunwoo Kim, Faeze Brahman +9
AI agents are increasingly autonomous in their interactions with human users and tools, leading to increased interactional safety risks. We present HAICOSYSTEM, a framework examini…
cs.AI2024
A Roadmap to Pluralistic Alignment
Taylor Sorensen, Jared Moore, Jillian Fisher +9
With increased power and prevalence of AI systems, it is ever more critical that AI systems are designed to serve all, i.e., people with diverse values and perspectives. However, a…