From the 1 of 3 linked papers with an AI index.
3 papers
Value Drifts: Tracing Value Alignment During LLM Post-Training
Mehar Bhatia, Shravan Nayak, Gaurav Kamath +4
The paper studies how large language models acquire and change their alignment with human values during post‑training, analyzing the impact of supervised fine‑tuning and preference…
On Optimizing Multimodal Jailbreaks for Spoken Language Models
Aravind Krishnan, Karolina StaÅczak, Dietrich Klakow
As Spoken Language Models (SLMs) integrate speech and text modalities, they inherit the safety vulnerabilities of their LLM backbone while introducing an expanded attack surface. S…
DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning
Sara Vera MarjanoviÄ, Arkil Patel, Vaibhav Adlakha +14
Large Reasoning Models like DeepSeek-R1 mark a fundamental shift in how LLMs approach complex problems. Instead of directly producing an answer for a given input, DeepSeek-R1 creat…