1 citations · 2 across the 27 of their papers we have counts for
4 papers · 1 filter
Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs
Zixuan Huang, Yang Zhou, Kaixuan Wang +7
Recent Vision-Language-Action (VLA) models for autonomous driving (AD) increasingly utilize chain-of-thought (CoT) supervision to enhance the reasoning capabilities of their Vision…
Adaptive Robust Estimator for Multi-Agent Reinforcement Learning
Zhongyi Li, Wan Tian, Jingyu Chen +8
Multi-agent collaboration has emerged as a powerful paradigm for enhancing the reasoning capabilities of large language models, yet it suffers from interaction-level ambiguity that…
CDRRM: Contrast-Driven Rubric Generation for Reliable and Interpretable Reward Modeling
Dengcan Liu, Fengkai Yang, Xiaohan Wang +7
Reward modeling is essential for aligning Large Language Models(LLMs) with human preferences, yet conventional reward models suffer from poor interpretability and heavy reliance on…
Federated Reasoning Distillation Framework with Model Learnability-Aware Data Allocation
Wei Guo, Siyuan Lu, Xiangdong Ran +8
Data allocation plays a critical role in federated large language model (LLM) and small language models (SLMs) reasoning collaboration. Nevertheless, existing data allocation metho…