Publications (18)
Emotion Concepts and their Function in a Large Language Model
Nicholas Sofroniew, Isaac Kauvar, William Saunders +13
Large language models (LLMs) sometimes appear to exhibit emotional reactions. We investigate why this is the case in Claude Sonnet 4.5 and explore implications for alignment-releva…
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Zhiwen Fan, Jian Zhang, Renjie Li +15
The rapid advancement of Large Multimodal Models (LMMs) for 2D images and videos has motivated extending these models to understand 3D scenes, aiming for human-like visual-spatial…
Neurosymbolic LoRA: Why and When to Tune Weights vs. Rewrite Prompts
Kevin Wang, Neel P. Bhatt, Cong Liu +7
Large language models (LLMs) can be adapted either through numerical updates that alter model parameters or symbolic manipulations that work on discrete prompts or logical constrai…
Extracting and Understanding the Superficial Knowledge in Alignment
Runjin Chen, Gabriel Jacob Perin, Xuxi Chen +5
Alignment of large language models (LLMs) with human values and preferences, often achieved through fine-tuning based on human feedback, is essential for ensuring safe and responsi…
Explaining Neural Networks Semantically and Quantitatively
Runjin Chen, Hao Chen, Ge Huang +2
This paper presents a method to explain the knowledge encoded in a convolutional neural network (CNN) quantitatively and semantically. The analysis of the specific rationale of eac…
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning
Gabriel J. Perin, Runjin Chen, Xuxi Chen +3
Large Language Models (LLMs) have become indispensable in real-world applications. However, their widespread adoption raises significant safety concerns, particularly in responding…