5 papers
Formalizing the Binding Problem
Lianghuan Huang, Yihao Li, Saeed Salehi +3
Representations of the world, arguably, contain information about features (e.g. something is blue, something is a circle) but also information about which features are part of the…
Causality Decodability, and Vice Versa: Lessons from Interpreting Counting ViTs
Lianghuan Huang, Yingshan Chang
Mechanistic interpretability seeks to uncover how internal components of neural networks give rise to predictions. A persistent challenge, however, is disentangling two often confl…
RAPID: An Efficient Reinforcement Learning Algorithm for Small Language Models
Lianghuan Huang, Sagnik Anupam, Insup Lee +2
Reinforcement learning (RL) has emerged as a promising strategy for finetuning small language models (SLMs) to solve targeted tasks such as math and coding. However, RL algorithms…
WST: Weak-to-Strong Knowledge Transfer via Reinforcement Learning
Haosen Ge, Shuo Li, Lianghuan Huang
Effective prompt engineering remains a challenging task for many applications. We introduce Weak-to-Strong Transfer (WST), an automatic prompt engineering framework where a small "…
Effective Reinforcement Learning for Reasoning in Language Models
Lianghuan Huang, Shuo Li, Sagnik Anupam +2
Reinforcement learning (RL) has emerged as a promising strategy for improving the reasoning capabilities of language models (LMs) in domains such as mathematics and coding. However…