1 citations · 1 across the 8 of their papers we have counts for
15 papers
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity
Anamika Lochab, Bolian Li, Ruqi Zhang
Reinforcement Learning with Verifiable Rewards (RLVR) has achieved substantial gains in single-attempt accuracy (Pass@1) on reasoning tasks, yet often suffers from reduced multi-sa…
Analytical Correction for Subsampling Bias in Drifting Models
Jiaru Zhang, Zeyun Deng, Juanwu Lu +2
Drifting models are capable one-step generative models trained to follow a drifting field. The field combines attractive and repulsive softmax-weighted centroids over the data and…
Learning From Developers: Towards Reliable Patch Validation at Scale for Linux
Chih-En Lin, Attreyee Mukherjee, Ajay Rawat +2
Patch reviewing is critical for software development, especially in distributed open-source development, which highly depends on voluntary work, such as Linux. This paper studies t…
Efficient and Explainable End-to-End Autonomous Driving via Masked Vision-Language-Action Diffusion
Jiaru Zhang, Manav Gagvani, Can Cui +3
Large Language Models (LLMs) and Vision-Language Models (VLMs) have emerged as promising candidates for end-to-end autonomous driving. However, these models typically face challeng…
Why Any-Order Autoregressive Models Need Two-Stream Attention: A Structural-Semantic Tradeoff
Patrick Pynadath, Ruqi Zhang
Any-order autoregressive models (AO-ARMs) offer a promising path toward efficient masked diffusion by enabling native key-value caching, but competitive performance has so far requ…
On Learning Closed-Loop Probabilistic Multi-Agent Simulator
Juanwu Lu, Rohit Gupta, Ahmadreza Moradipari +3
The rapid iteration of autonomous vehicle (AV) deployments leads to increasing needs for building realistic and scalable multi-agent traffic simulators for efficient evaluation. Re…