Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
HEED: Density-Weighted Residual Alignment for Hybrid Vision-Language Model Distillation
Yihao Liang, Niraj K. Jha
Distilling vision-language models into faster hybrid architectures, such as 3:1 Mamba-2/attention mixes, is now standard practice for making inference efficient. Aggregate benchmar…
cs.CV2026
DVD: Deterministic Video Depth Estimation with Generative Priors
Hongfei Zhang, Harold Haodong Chen, Chenfei Liao +12
Existing video depth estimation faces a fundamental trade-off: generative models suffer from stochastic geometric hallucinations and scale drift, while discriminative models demand…