9 papers
PrAg-PO: Prompt Augmented Policy Optimization for Robust and Diverse Mathematical Reasoning
Wenquan Lu, Hai Huang, Enqi Liu +1
Reinforcement learning algorithms such as group-relative policy optimization (GRPO) have shown strong potential for improving the mathematical reasoning capabilities of large langu…
Opinion: Towards Unified Expressive Policy Optimization for Robust Robot Learning
Haidong Huang, Haiyue Zhu. Jiayu Song, Xixin Zhao +4
Offline-to-online reinforcement learning (O2O-RL) has emerged as a promising paradigm for safe and efficient robotic policy deployment but suffers from two fundamental challenges:…
UMDAM: A Unified Data Layout and DRAM Address Mapping for Heterogenous NPU-PIM
Hai Huang
Large Language Models (LLMs) are increasingly deployed on edge devices with Neural Processing Units (NPUs), yet the decode phase remains memory-intensive, limiting performance. Pro…
MSM-Seg: A Modality-and-Slice Memory Framework with Category-Agnostic Prompting for Multi-Modal Brain Tumor Segmentation
Yuxiang Luo, Qing Xu, Hai Huang +3
Multi-modal brain tumor segmentation is critical for clinical diagnosis, and it requires accurate identification of distinct internal anatomical subregions. While the recent prompt…
LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
Hai Huang, Yann LeCun, Randall Balestriero
Large Language Model (LLM) pretraining, finetuning, and evaluation rely on input-space reconstruction and generative capabilities. Yet, it has been observed in vision that embeddin…
LoRA Users Beware: A Few Spurious Tokens Can Manipulate Your Finetuned Model
Marcel Mateos Salles, Praney Goyal, Pradyut Sekhsaria +2
Large Language Models (LLMs) are commonly finetuned for a variety of use cases and domains. A common approach is to leverage Low-Rank Adaptation (LoRA) -- known to provide strong p…