4 papers
Localizing RL-Induced Tool Use to a Single Crosscoder Feature
Andrii Shportko, Shubham Bhokare, Ahmed Zeyad A Alzahrani +3
Fine-tuning through RL reshapes the internal representations of language models to enable agentic behaviors such as tool use, yet the mechanistic basis of these changes remains poo…
When Choices Become Risks: Safety Failures of Large Language Models under Multiple-Choice Constraints
Yuheng Chen, Zhiyu Wu, Bowen Cheng +1
Safety alignment in large language models (LLMs) is primarily evaluated under open-ended generation, where models can mitigate risk by refusing to respond. In contrast, many real-w…
Balancing Knowledge Delivery and Emotional Comfort in Healthcare Conversational Systems
Shang-Chi Tsai, Yun-Nung Chen
With the advancement of large language models, many dialogue systems are now capable of providing reasonable and informative responses to patients' medical conditions. However, whe…
GPT-4o System Card
OpenAI, :, Aaron Hurst +416
GPT-4o is an autoregressive omni model that accepts as input any combination of text, audio, image, and video, and generates any combination of text, audio, and image outputs. It's…