most citedDefining and Evaluating Physical Safety for Large Language Models

1 citations · 1 across the 3 of their papers we have counts for

collaborators

5 papers

cs.AI2026

Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT

Cen Lu, Yung-Chen Tang, Andrea Cavallaro

Developers judge a model checkpoint by how it behaves. After supervised fine-tuning (SFT), two checkpoints that perform about the same across relevant benchmarks are treated as int…

cs.AI2026

Sparse Neuron Ablation Triggers Catastrophic Collapse of the Language Core in Large Vision-Language Models

Cen Lu, Yung-Chen Tang, Andrea Cavallaro

Large Vision-Language Models (LVLMs) have shown impressive multimodal understanding capabilities, yet the structures that sustain their functionality remain poorly understood from…

cs.LG20261 cited

Defining and Evaluating Physical Safety for Large Language Models

Yung-Chen Tang, Pin-Yu Chen, Tsung-Yi Ho

Large Language Models (LLMs) are increasingly used to control robotic systems such as drones, but their risks of causing physical threats and harm in real-world applications remain…

cs.LG2025

CarBoN: Calibrated Best-of-N Sampling Improves Test-time Reasoning

Yung-Chen Tang, Pin-Yu Chen, Andrea Cavallaro

Allocating more computation during inference time (test-time scaling) improves language model performance, especially for reasoning tasks. However, popular methods like Best-of-

cs.CR2025

Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets

Lei Hsiung, Tianyu Pang, Yung-Chen Tang +4

Recent advancements in large language models (LLMs) have underscored their vulnerability to safety alignment jailbreaks, particularly when subjected to downstream fine-tuning. Howe…