4 papers
When Autoregressive Consistency Hurts Safety Alignment
Bochen Lyu, Yiyang Jia, Xiaohao Cai +1
Safety alignment in large language models (LLMs) is fragile in part because it is often shallow: fine-tuning mainly reshapes the model's behavior near the first few output tokens.…
Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently
Bochen Lyu, Yiyang Jia, Xiaohao Cai +1
Transformers can acquire Chain-of-Thought (CoT) capabilities to solve reasoning tasks via fine-tuning. Reinforcement learning (RL) and supervised fine-tuning (SFT) are two primary…
SEIS: Subspace-based Equivariance and Invariance Scores for Neural Representations
Huahua Lin, Katayoun Farrahi, Xiaohao Cai
Understanding how neural representations respond to geometric transformations is essential for evaluating whether learned features preserve meaningful spatial structure. Existing a…
From Instance Segmentation to 3D Growth Trajectory Reconstruction in Planktonic Foraminifera
Huahua Lin, Xiaohao Cai, Mark Nixon +2
Planktonic foraminifera, marine protists characterized by their intricate chambered shells, serve as valuable indicators of past and present environmental conditions. Understanding…