1 citations · 1 across the 3 of their papers we have counts for
5 papers
Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT
Cen Lu, Yung-Chen Tang, Andrea Cavallaro
Developers judge a model checkpoint by how it behaves. After supervised fine-tuning (SFT), two checkpoints that perform about the same across relevant benchmarks are treated as int…
Sparse Neuron Ablation Triggers Catastrophic Collapse of the Language Core in Large Vision-Language Models
Cen Lu, Yung-Chen Tang, Andrea Cavallaro
Large Vision-Language Models (LVLMs) have shown impressive multimodal understanding capabilities, yet the structures that sustain their functionality remain poorly understood from…
Defining and Evaluating Physical Safety for Large Language Models
Yung-Chen Tang, Pin-Yu Chen, Tsung-Yi Ho
Large Language Models (LLMs) are increasingly used to control robotic systems such as drones, but their risks of causing physical threats and harm in real-world applications remain…
CarBoN: Calibrated Best-of-N Sampling Improves Test-time Reasoning
Yung-Chen Tang, Pin-Yu Chen, Andrea Cavallaro
Allocating more computation during inference time (test-time scaling) improves language model performance, especially for reasoning tasks. However, popular methods like Best-of-…
Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
Lei Hsiung, Tianyu Pang, Yung-Chen Tang +4
Recent advancements in large language models (LLMs) have underscored their vulnerability to safety alignment jailbreaks, particularly when subjected to downstream fine-tuning. Howe…