12 papers
Revealing Safety-Critical Scenarios for UTM via Transformer
Huaze Tang, Bill Zeng, Chao Wang +3
Unmanned Traffic Management (UTM) systems are cloud-based platforms designed to manage and coordinate multiple aerial vehicles remotely. UTM systems are safety-critical which canno…
Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners
Chao Wang, Hongtao Tian, Tao Yang +3
Group Relative Policy Optimization (GRPO) is a default recipe for process-supervised reinforcement learning of LLM reasoners, and dense process supervision -- via learned process r…
dVLA-RL: Reinforcement Learning over Denoising Trajectories for Discrete Diffusion Vision-Language-Action Models
Yuhao Wu, Yitian Liu, Weijie Shen +13
Vision-Language-Action (VLA) models have established a powerful paradigm for generalist robotic manipulation by grounding control into the semantic reasoning of VLMs. Prevailing ar…
Lightweight and Generalizable Multi-Sensor Human Activity Recognition via Cascaded Fusion and Style-Augmented Decomposition
Wang Chenglong, Zhuo Yan, Ding Wenbo +1
Wearable Human Activity Recognition (WHAR) is a prominent research area within ubiquitous computing, whose core lies in effectively modeling intra- and inter-sensor spatio-temporal…
RUMAD: Reinforcement-Unifying Multi-Agent Debate
Chao Wang, Han Lin, Huaze Tang +2
Multi-agent debate (MAD) systems leverage collective intelligence to enhance reasoning capabilities, yet existing approaches struggle to simultaneously optimize accuracy, consensus…
ProAgentBench: Evaluating LLM Agents for Proactive Assistance with Real-World Data
Yuanbo Tang, Huaze Tang, Tingyu Cao +6
Proactive agents that anticipate user intentions without explicit prompts represent a significant evolution in human-AI interaction, promising to reduce cognitive load and streamli…