5 papers
Opinion: Towards Unified Expressive Policy Optimization for Robust Robot Learning
Haidong Huang, Haiyue Zhu. Jiayu Song, Xixin Zhao +4
Offline-to-online reinforcement learning (O2O-RL) has emerged as a promising paradigm for safe and efficient robotic policy deployment but suffers from two fundamental challenges:…
LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
Hai Huang, Yann LeCun, Randall Balestriero
Large Language Model (LLM) pretraining, finetuning, and evaluation rely on input-space reconstruction and generative capabilities. Yet, it has been observed in vision that embeddin…
Directed Information -covering: An Information-Theoretic Framework for Context Engineering
Hai Huang
We introduce \textbf{Directed Information -covering}, a simple but general framework for redundancy-aware context engineering. Directed information (DI), a causal analogue of mu…
LoRA Users Beware: A Few Spurious Tokens Can Manipulate Your Finetuned Model
Marcel Mateos Salles, Praney Goyal, Pradyut Sekhsaria +2
Large Language Models (LLMs) are commonly finetuned for a variety of use cases and domains. A common approach is to leverage Low-Rank Adaptation (LoRA) -- known to provide strong p…
Next Token Perception Score: Analytical Assessment of your LLM Perception Skills
Yu-Ang Cheng, Leyang Hu, Hai Huang +1
Autoregressive pretraining has become the de facto paradigm for learning general-purpose representations in large language models (LLMs). However, linear probe performance across d…