1 citations · 1 across the 5 of their papers we have counts for
9 papers
Large Reasoning Models Learn Better Alignment from Flawed Thinking
ShengYun Peng, Pin-Yu Chen, Eric Smith +6
Large reasoning models (LRMs) "think" by generating structured chain-of-thought (CoT) before producing a final answer, yet they still lack the ability to reason critically about sa…
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety
Seongmin Lee, Aeree Cho, Grace C. Kim +3
As large language models (LLMs) see wider real-world use, understanding and mitigating their unsafe behaviors is critical. Interpretation techniques can reveal causes of unsafe out…
Shape it Up! Restoring LLM Safety during Finetuning
ShengYun Peng, Pin-Yu Chen, Jianfeng Chi +2
Finetuning large language models (LLMs) enables user-specific customization but introduces critical safety risks: even a few harmful examples can compromise safety alignment. A com…
Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language Models
ShengYun Peng, Pin-Yu Chen, Matthew Hull +1
Safety alignment is crucial to ensure that large language models (LLMs) behave in ways that align with human preferences and prevent harmful actions during inference. However, rece…
Interactive Visual Learning for Stable Diffusion
Seongmin Lee, Benjamin Hoover, Hendrik Strobelt +7
Diffusion-based generative models' impressive ability to create convincing images has garnered global attention. However, their complex internal structures and operations often pos…
LLM Attributor: Interactive Visual Attribution for LLM Generation
Seongmin Lee, Zijie J. Wang, Aishwarya Chakravarthy +5
While large language models (LLMs) have shown remarkable capability to generate convincing text across diverse domains, concerns around its potential risks have highlighted the imp…