3 papers
cs.CV2026
DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models
Zhuoming Liu, Jinhong Lin, Kwan Man Cheng +3
Many modern vision-language models (VLMs) build on autoregressive decoding of discrete tokens. While text-based output interfaces enable scalable pretraining and strong zero-shot g…
cs.CV2026
An Extended VIIRS-like Artificial Nighttime Light Data Reconstruction (1986-2024)
Yihe Tian, Kwan Man Cheng, Zhengbo Zhang +6
Artificial Night-Time Light (NTL) remote sensing is a vital proxy for quantifying the intensity and spatial distribution of human activities. Although the NPP-VIIRS sensor provides…
cs.CV2025
Modular-Cam: Modular Dynamic Camera-view Video Generation with LLM
Zirui Pan, Xin Wang, Yipeng Zhang +4
Text-to-Video generation, which utilizes the provided text prompt to generate high-quality videos, has drawn increasing attention and achieved great success due to the development…