From the 1 of 21 linked papers with an AI index.
21 papers
Privileged Likelihood Is Not Automatically Value: Three Checks for Token Credit in On-Policy Self-Distillation
Xuan-Phi Nguyen, Shrey Pandit, Yiran Zhao +3
Outcome verifiers score completed reasoning traces but do not assign credit to intermediate tokens. Privileged self-distillation attempts to fill this gap by rescoring a model's ow…
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models
Yidong Wang, Yan Zhan, Ziteng Feng +16
Reward models are a bottleneck for reinforcement learning in embodied AI. Long-horizon robotic manipulation requires scalable vision feedback beyond handcrafted rewards or task-spe…
Mental World Modeling
Hao Fei, Yiran Zhao
The paper introduces Mental World Modeling (MWM), a framework that integrates agents' mental states (beliefs, desires, intentions) into world models, and presents a training‑free b…
Mixture-of-Parallelisms: Towards Memory-Efficient Training Stack for Mixture-of-Experts Models
Xuan-Phi Nguyen, Shrey Pandit, Yiran Zhao +3
This paper showcases a memory-efficient training stack for Mixture-of-Experts (MoE) models. It is a training paradigm that combines and specializes various existing and novel paral…
Multilingual Fine-Tuning via Localized Gradient Conflict Resolution
Long P. Hoang, Yiran Zhao, Wei Lu +1
The rapid evolution of Large Language Models (LLMs) has established cross-lingual versatility as a defining feature of modern systems. However, fine-tuning these models frequently…
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments
Chiyu Zhang, Huiqin Yang, Bendong Jiang +8
The rapid proliferation of LLM-based autonomous agents in real operating system environments introduces a new category of safety risk beyond content safety: behavior jailbreak, whe…