10 citations · 15 across the 11 of their papers we have counts for
3 papers · 1 filter
ODESteer: A Unified ODE-Based Steering Framework for LLM Alignment
Hongjue Zhao, Haosen Sun, Jiangtao Kong +8
Activation steering, or representation engineering, offers a lightweight approach to align large language models (LLMs) by manipulating their internal activations at inference time…
A Roadmap to Pluralistic Alignment
Taylor Sorensen, Jared Moore, Jillian Fisher +9
With increased power and prevalence of AI systems, it is ever more critical that AI systems are designed to serve all, i.e., people with diverse values and perspectives. However, a…
The Generative AI Paradox: "What It Can Create, It May Not Understand"
Peter West, Ximing Lu, Nouha Dziri +11
The recent wave of generative AI has sparked unprecedented global attention, with both excitement and concern over potentially superhuman levels of artificial intelligence: models…