activity
20152026
most citedJoint Stereo Video Deblurring, Scene Flow Estimation and Moving Object Segmentation

29 citations · 158 across the 59 of their papers we have counts for

collaborators

89 papers

cs.RO2026

CLEAR: Closed-Loop Reinforcement Learning at Scale for End-to-End Autonomous Driving

Yunxiao Shi, Hong Cai, Mohammad Ghavamzadeh +1

End-to-end autonomous driving (E2E-AD) aims to directly map raw sensor information to driving actions. Recently, with the rapid advancement of multi-modal large language models (ML…

cs.CV2026

SciFlow: Semantic Cross Interference for Self-Supervised Optical Flow Domain Generalization

Jamie Menjay Lin, Jisoo Jeong, Hong Cai +2

Motions of objects and scenes carry essential intelligence in video understanding, offering rich cues for interpreting dynamic settings and interactions. Due to the cost and scarci…

cs.CV2025

Neodragon: Mobile Video Generation using Diffusion Transformer

Animesh Karnewar, Denis Korzhenkov, Ioannis Lelekas +10

We introduce Neodragon, a text-to-video system capable of generating 2s (49 frames @24 fps) videos at the 640x1024 resolution directly on a Qualcomm Hexagon NPU in a record 6.7s (7…

cs.CV2025

ODG: Occupancy Prediction Using Dual Gaussians

Yunxiao Shi, Yinhao Zhu, Shizhong Han +4

Occupancy prediction infers fine-grained 3D geometry and semantics from camera images of the surrounding environment, making it a critical perception task for autonomous driving. E…

cs.CV2025

DySS: Dynamic Queries and State-Space Learning for Efficient 3D Object Detection from Multi-Camera Videos

Rajeev Yasarla, Shizhong Han, Hong Cai +1

Camera-based 3D object detection in Bird's Eye View (BEV) is one of the most important perception tasks in autonomous driving. Earlier methods rely on dense BEV features, which are…

cs.CV2025

Learning Optical Flow Field via Neural Ordinary Differential Equation

Leyla Mirvakhabova, Hong Cai, Jisoo Jeong +3

Recent works on optical flow estimation use neural networks to predict the flow field that maps positions of one image to positions of the other. These networks consist of a featur…