1 citations · 1 across the 3 of their papers we have counts for
20 papers
Average-Power-Budgeted Underwater Vehicle Control via Constrained Reinforcement Learning
Yinuo Wang, Gavin Tao, Yuze Liu +1
Underwater vehicles operate from a fixed onboard energy budget that propulsion rapidly depletes, so a controller that completes its task while drawing less thruster power directly…
OmniGAIA: Towards Native Omni-Modal AI Agents
Xiaoxi Li, Wenxiang Jiao, Jiarui Jin +10
Human intelligence naturally intertwines omni-modal perception -- spanning vision, audio, and language -- with complex reasoning and tool usage to interact with the world. However,…
Factor-Aware Mixture-of-Experts with Pretrained Encoder for Combinatorial Generalization
Feihong Zhang, Guojian Zhan, Zeyu He +8
The integration of pretrained encoders with diffusion policies has become a dominant paradigm for visual robotic manipulation. However, it still struggles to generalize across comp…
STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens
Shiqi Liu, Zeyu He, Guojian Zhan +10
Reinforcement Learning (RL) has significantly improved large language model reasoning, but existing RL fine-tuning methods rely heavily on heuristic techniques such as entropy regu…
Optimal Transport for LLM Reward Modeling from Noisy Preference
Licheng Pan, Haochen Yang, Haoxuan Li +8
Reward models are fundamental to Reinforcement Learning from Human Feedback (RLHF), yet real-world datasets are inevitably corrupted by noisy preference. Conventional training obje…
ImplicitRM: Unbiased Reward Modeling from Implicit Preference Data for LLM alignment
Hao Wang, Haocheng Yang, Licheng Pan +7
Reward modeling represents a long-standing challenge in reinforcement learning from human feedback (RLHF) for aligning language models. Current reward modeling is heavily contingen…