papers

Publications (35)

cond-mat.mtrl-sci2019

Planar topological Hall effect in a uniaxial van der Waals ferromagnet Fe3GeTe2

Yurong You, Yuanyuan Gong, Hang Li +8

In this work, we reported the observation of a novel planar topological Hall effect (PTHE) in single crystal of Fe3GeTe2, a paradigmatic two-dimensional ferromagnet with strong uni…

cs.LG2019

Simple Black-box Adversarial Attacks

Chuan Guo, Jacob R. Gardner, Yurong You +2

We propose an intriguingly simple method for the construction of adversarial images in the black-box setting. In constrast to the white-box scenario, constructing black-box adversa…

cs.AI2017

Virtual to Real Reinforcement Learning for Autonomous Driving

Xinlei Pan, Yurong You, Ziyan Wang +1

Reinforcement learning is considered as a promising direction for driving policy learning. However, training autonomous driving vehicle with reinforcement learning in real environm…

cs.CV2024

DiffuBox: Refining 3D Object Detection with Point Diffusion

Xiangyu Chen, Zhenzhen Liu, Katie Z Luo +10

Ensuring robust 3D object detection and localization is crucial for many applications in robotics and autonomous driving. Recent models, however, face difficulties in maintaining h…

cs.CV2024

Pre-Training LiDAR-Based 3D Object Detectors Through Colorization

Tai-Yu Pan, Chenyang Ma, Tianle Chen +7

Accurate 3D object detection and understanding for self-driving cars heavily relies on LiDAR point clouds, necessitating large amounts of labeled data to train. In this work, we in…

cond-mat.mtrl-sci2016

Designing compensated magnetic states in tetragonal Mn3Ge-based alloys

Yurong You, Guizhou Xu, Fang Hu +3

Magnetic compensated state attracted much interests due to the observed large exchange bias and large coercivity, and its potential applications in the antiferromagnetic spintronic…

cs.RO2026

Planning-aligned Token Compression for Long-Context Autonomous Driving

Zhixuan Liang, Yuxiao Chen, Yurong You +12

Monolithic vision-action models represent an emerging paradigm in autonomous driving. However, this architecture produces token sequences that quickly exceed real-time computationa…

cs.CV2023

Unsupervised Domain Adaptation for Self-Driving from Past Traversal Features

Travis Zhang, Katie Luo, Cheng Perng Phoo +5

The rapid development of 3D object detection systems for self-driving cars has significantly improved accuracy. However, these systems struggle to generalize across diverse driving…

cs.CV2022

Exploiting Playbacks in Unsupervised Domain Adaptation for 3D Object Detection

Yurong You, Carlos Andres Diaz-Ruiz, Yan Wang +4

Self-driving cars must detect other vehicles and pedestrians in 3D to plan safe routes and avoid collisions. State-of-the-art 3D object detectors, based on deep learning, have show…

cond-mat.mtrl-sci2019

Design of reversible low-field magnetocaloric effect at room temperature in hexagonal MnMX ferromagnets

Jun Liu, Yurong You, Ivan Batashev +8

Giant magnetocaloric effect is widely achieved in hexagonal MnMX-based (M = Co or Ni, X = Si or Ge) ferromagnets at their first-order magnetostructural transition. However, the the…

cs.CV2026

Cosmos 3: Omnimodal World Models for Physical AI

NVIDIA, :, Aditi +293

We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…

cs.CV2020

Train in Germany, Test in The USA: Making 3D Object Detectors Generalize

Yan Wang, Xiangyu Chen, Yurong You +5

In the domain of autonomous driving, deep learning has substantially improved the 3D object detection accuracy for LiDAR and stereo camera data alike. While deep networks are great…

cs.RO2026

Accelerating Structured Chain-of-Thought in Autonomous Vehicles

Yi Gu, Yan Wang, Yuxiao Chen +8

Chain-of-Thought (CoT) reasoning enhances the decision-making capabilities of vision-language-action models in autonomous driving, but its autoregressive nature introduces signific…

cs.CV2024

STORM: Spatio-Temporal Reconstruction Model for Large-Scale Outdoor Scenes

Jiawei Yang, Jiahui Huang, Yuxiao Chen +10

We present STORM, a spatio-temporal reconstruction model designed for reconstructing dynamic outdoor scenes from sparse observations. Existing dynamic reconstruction methods often…

cond-mat.mtrl-sci2018

Tunable magnetic and transport properties of Mn3Ga thin films on Ta/Ru seedlayer

Fang Hu, Guizhou Xu, Yurong You +8

Hexagonal D019-type Mn3Z alloys that possess large anomalous and topological-like Hall effects have attracted much attention due to their great potential in the antiferromagnetic s…

cs.CV2020

End-to-End Pseudo-LiDAR for Image-Based 3D Object Detection

Rui Qian, Divyansh Garg, Yan Wang +6

Reliable and accurate 3D object detection is a necessity for safe autonomous driving. Although LiDAR sensors can provide accurate 3D point cloud estimates of the environment, they…

cs.CV2026

Latent Chain-of-Thought World Modeling for End-to-End Driving

Shuhan Tan, Kashyap Chitta, Yuxiao Chen +8

Recent Vision-Language-Action (VLA) models for autonomous driving explore inference-time reasoning as a way to improve driving performance and safety in challenging scenarios. Most…

cs.CV2025

Towards Efficient and Effective Multi-Camera Encoding for End-to-End Driving

Jiawei Yang, Ziyu Chen, Yurong You +7

We present Flex, an efficient and effective scene encoder that addresses the computational bottleneck of processing high-volume multi-camera data in end-to-end autonomous driving.…

cs.CV2022

Depth Estimation Matters Most: Improving Per-Object Depth Estimation for Monocular 3D Detection and Tracking

Longlong Jing, Ruichi Yu, Henrik Kretzschmar +11

Monocular image-based 3D perception has become an active research area in recent years owing to its applications in autonomous driving. Approaches to monocular 3D perception includ…

cs.RO2026

Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail

NVIDIA, :, Yan Wang +41

End-to-end architectures trained via imitation learning have advanced autonomous driving by scaling model size and data, yet performance remains brittle in safety-critical long-tai…

cs.CV2024

Better Monocular 3D Detectors with LiDAR from the Past

Yurong You, Cheng Perng Phoo, Carlos Andres Diaz-Ruiz +5

Accurate 3D object detection is crucial to autonomous driving. Though LiDAR-based detectors have achieved impressive performance, the high cost of LiDAR sensors precludes their wid…

cs.CV2020

Pseudo-LiDAR++: Accurate Depth for 3D Object Detection in Autonomous Driving

Yurong You, Yan Wang, Wei-Lun Chao +5

Detecting objects such as cars and pedestrians in 3D plays an indispensable role in autonomous driving. Existing approaches largely rely on expensive LiDAR sensors for accurate dep…

cs.CV2025

Extrapolated Urban View Synthesis Benchmark

Xiangyu Han, Zhen Jia, Boyi Li +8

Photorealistic simulators are essential for the training and evaluation of vision-centric autonomous vehicles (AVs). At their core is Novel View Synthesis (NVS), a crucial capabili…

cs.CV2022

R4D: Utilizing Reference Objects for Long-Range Distance Estimation

Yingwei Li, Tiffany Chen, Maya Kabkab +4

Estimating the distance of objects is a safety-critical task for autonomous driving. Focusing on short-range objects, existing methods and datasets neglect the equally important lo…

cs.CV2022

Ithaca365: Dataset and Driving Perception under Repeated and Challenging Weather Conditions

Carlos A. Diaz-Ruiz, Youya Xia, Yurong You +11

Advances in perception for self-driving cars have accelerated in recent years due to the availability of large-scale datasets, typically collected at specific locations and under n…

cs.CV2022

Hindsight is 20/20: Leveraging Past Traversals to Aid 3D Perception

Yurong You, Katie Z Luo, Xiangyu Chen +6

Self-driving cars must detect vehicles, pedestrians, and other traffic participants accurately to operate safely. Small, far-away, or highly occluded objects are particularly chall…

cs.CV2018

Resource Aware Person Re-identification across Multiple Resolutions

Yan Wang, Lequn Wang, Yurong You +6

Not all people are equally easy to identify: color statistics might be enough for some cases while others might require careful reasoning about high- and low-level details. However…

cs.CV2025

DreamDrive: Generative 4D Scene Modeling from Street View Images

Jiageng Mao, Boyi Li, Boris Ivanovic +7

Synthesizing photo-realistic visual observations from an ego vehicle's driving trajectory is a critical step towards scalable training of self-driving models. Reconstruction-based…

cs.CV2023

Unsupervised Adaptation from Repeated Traversals for Autonomous Driving

Yurong You, Cheng Perng Phoo, Katie Z Luo +5

For a self-driving car to operate reliably, its perceptual system must generalize to the end-user's environment -- ideally without additional annotation efforts. One potential solu…

cond-mat.mtrl-sci2023

Growth of high-quality CrI3 single crystals and engineering of its magnetic properties via V and Mn doping

Shuang Pan, Yuqing Bai, Jiaxuan Tang +4

CrI3, as a soft van der Waals layered magnetic material, has been widely concerned and explored for its magnetic complexity and tunability. In this work, high quality and large siz…

cs.CV2022

Learning to Detect Mobile Objects from LiDAR Scans Without Labels

Yurong You, Katie Z Luo, Cheng Perng Phoo +5

Current 3D object detectors for autonomous driving are almost entirely trained on human-annotated data. Although of high quality, the generation of such data is laborious and costl…

cs.CV2024

Language-Image Models with 3D Understanding

Jang Hyun Cho, Boris Ivanovic, Yulong Cao +8

Multi-modal large language models (MLLMs) have shown incredible capabilities in a variety of 2D vision and language tasks. We extend MLLMs' perceptual capabilities to ground and re…

cs.RO2025

Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive Reasoning

Zhenghao "Mark" Peng, Wenhao Ding, Yurong You +11

Recent reasoning-augmented Vision-Language-Action (VLA) models have improved the interpretability of end-to-end autonomous driving by generating intermediate reasoning traces. Yet…

cs.CV2023

Reward Finetuning for Faster and More Accurate Unsupervised Object Discovery

Katie Z Luo, Zhenzhen Liu, Xiangyu Chen +7

Recent advances in machine learning have shown that Reinforcement Learning from Human Feedback (RLHF) can improve machine learning models and align them with human preferences. Alt…

cs.CV2025

Efficient Multi-Camera Tokenization with Triplanes for End-to-End Driving

Boris Ivanovic, Cristiano Saltori, Yurong You +3

Autoregressive Transformers are increasingly being deployed as end-to-end robot and autonomous vehicle (AV) policy architectures, owing to their scalability and potential to levera…