26 papers
Anchored, Not Graded: Vision-Language Models Fail at Slant-from-Texture Perception
Qian Zhang, Michal Golovanevsky, Fulvio Domini +1
Human perception of surface slant from texture exhibits systematic, graded biases that emerge reliably in psychophysical experiments. Prior work showed that unsupervised CNNs repro…
Anchor3R: Streaming 3D Reconstruction with Transient Anchors for Long-Horizon Visual Mapping
Peilin Tao, Chong Cheng, Yuansen Du +8
Long-horizon online visual mapping is a core capability for robot perception, requiring continuous camera-motion and scene-geometry estimation from visual streams under bounded mem…
HorizonStream: Long-Horizon Attention for Streaming 3D Reconstruction
Chong Cheng, Peilin Tao, Nanjie Yao +9
Online 3D reconstruction requires estimating camera pose and scene geometry under strict causal and bounded-memory constraints. Existing methods often suffer from drift, jitter, or…
Operator Learning for Reconstructing Flow Fields from Sparse Measurements: a Language Model Approach
Qian Zhang, George Em Karniadakis
Reconstructing flow fields from sparse measurements is a fundamental problem in fluid mechanics with broad implications for modeling, control, and design. In this work, we propose…
HorizonDrive: Self-Corrective Autoregressive World Model for Long-horizon Driving Simulation
Conglang Zhang, Yifan Zhan, Qingjie Wang +10
Closed-loop driving simulation requires real-time interaction beyond short offline clips, pushing current driving world models toward autoregressive (AR) rollout. Existing AR disti…
EponaV2: Driving World Model with Comprehensive Future Reasoning
Jiawei Xu, Zhizhou Zhong, Zhijian Shu +8
Data scaling plays a pivotal role in the pursuit of general intelligence. However, the prevailing perception-planning paradigm in autonomous driving relies heavily on expensive man…