5 papers
QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization
Siwei Chen, Siqi Chen, Xupeng Miao +1
Recent large reasoning models often develop long chain-of-thought responses during reinforcement learning (RL), resulting in high inference latency and deployment cost. Existing me…
DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning
Yujie Wang, Siwei Chen, Longzan Luo +4
Reinforcement Learning (RL) has become pivotal for improving model capabilities yet suffers from rollout efficiency bottlenecks due to the long-tail response length distribution. W…
Conditionally Whitened Generative Models for Probabilistic Time Series Forecasting
Yanfeng Yang, Siwei Chen, Pingping Hu +6
Probabilistic forecasting of multivariate time series is challenging due to non-stationarity, inter-variable dependencies, and distribution shifts. While recent diffusion and flow…
WHALES: A Multi-Agent Scheduling Dataset for Enhanced Cooperation in Autonomous Driving
Yinsong Wang, Siwei Chen, Ziyi Song +1
Cooperative perception research is hindered by the limited availability of datasets that capture the complexity of real-world Vehicle-to-Everything (V2X) interactions, particularly…
Do Graph Diffusion Models Accurately Capture and Generate Substructure Distributions?
Xiyuan Wang, Yewei Liu, Lexi Pang +2
Diffusion models have gained popularity in graph generation tasks; however, the extent of their expressivity concerning the graph distributions they can learn is not fully understo…