7 papers
FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving
Zekai Li, Yihao Liang, Hongfei Zhang +3
Vision-Language-Action (VLA) models promise to bring end-to-end reasoning to autonomous driving, but their computational cost remains far too high for real-time control. The core c…
An Alternative Trajectory for Generative AI
Margarita Belova, Yuval Kansal, Yihao Liang +2
The generative artificial intelligence (AI) ecosystem is undergoing rapid transformations that threaten its sustainability. As models transition from research prototypes to high-tr…
HEED: Density-Weighted Residual Alignment for Hybrid Vision-Language Model Distillation
Yihao Liang, Niraj K. Jha
Distilling vision-language models into faster hybrid architectures, such as 3:1 Mamba-2/attention mixes, is now standard practice for making inference efficient. Aggregate benchmar…
DVD: Deterministic Video Depth Estimation with Generative Priors
Hongfei Zhang, Harold Haodong Chen, Chenfei Liao +12
Existing video depth estimation faces a fundamental trade-off: generative models suffer from stochastic geometric hallucinations and scale drift, while discriminative models demand…
CD4LM: Consistency Distillation and aDaptive Decoding for Diffusion Language Models
Yihao Liang, Ze Wang, Hao Chen +7
Autoregressive large language models achieve strong results on many benchmarks, but decoding remains fundamentally latency-limited by sequential dependence on previously generated…
The Universal Landscape of Human Reasoning
Qiguang Chen, Jinhao Liu, Libo Qin +14
Understanding how information is dynamically accumulated and transformed in human reasoning has long challenged cognitive psychology, philosophy, and artificial intelligence. Exist…