4 papers
Synchronized Logit Steering: Real-world Steganography
Andrew Rufail, Aadi Dash, Onir Narahari +5
Steganography in large language models offers a way to embed hidden messages within natural-sounding text. Existing token and logit-level methods typically require the sender and r…
Semantic Convergence: Investigating Shared Representations Across Scaled LLMs
Daniel Son, Sanjana Rathore, Andrew Rufail +6
We investigate feature universality in Gemma-2 language models (Gemma-2-2B and Gemma-2-9B), asking whether models with a four-fold difference in scale still converge on comparable…
Introducing MAPO: Momentum-Aided Gradient Descent Prompt Optimization
Anthony Cui, Pranav Nandyalam, Andrew Rufail +4
Momentum-Aided Prompt Optimization (MAPO) enhances the efficiency and efficacy of prompt optimization for Large Language Models (LLMs). Building on ProTeGi, MAPO uses positive natu…
CLEAR: Contrasting Textual Feedback with Experts and Amateurs for Reasoning
Andrew Rufail, Daniel Kim, Sean O'Brien +1
We introduce CLEAR (Contrasting Textual Feedback with Experts and Amateurs for Reasoning), a novel approach to language model reasoning that leverages the strengths of a larger (ex…