5 papers
When Data Imbalance Helps: Robust Generalization Through Shortcut Saturation
Cheng-Ting Chou, Duc Binh Hoang
We study robust generalization under spurious correlations: tasks where a shortcut feature is correlated with the true label in training but anti-correlated in an adversarial held-…
CIPHER: Cryptographic Insecurity Profiling via Hybrid Evaluation of Responses
Max Manolov, Tony Gao, Siddharth Shukla +2
Large language models (LLMs) are increasingly used to assist developers with code, yet their implementations of cryptographic functionality often contain exploitable flaws. Minor d…
CLaS-Bench: A Cross-Lingual Alignment and Steering Benchmark
Daniil Gurgurov, Yusser Al Ghussin, Tanja Baeumel +5
Understanding and controlling the behavior of large language models (LLMs) is an increasingly important topic in multilingual NLP. Beyond prompting or fine-tuning, , i.e.,~manipula…
Universal Neurons in GPT-2: Emergence, Persistence, and Functional Impact
Advey Nandan, Cheng-Ting Chou, Amrit Kurakula +4
We investigate the phenomenon of neuron universality in independently trained GPT-2 Small models, examining these universal neurons-neurons with consistently correlated activations…
Causal Language Control in Multilingual Transformers via Sparse Feature Steering
Cheng-Ting Chou, George Liu, Jessica Sun +4
Deterministically controlling the target generation language of large multilingual language models (LLMs) remains a fundamental challenge, particularly in zero-shot settings where…