3 papers
cs.CV2026
ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding
Mingkang Dong, Muxin Pu, Jie Li +8
Streaming video understanding requires models to continuously retain useful visual evidence before future questions are known. Existing approaches primarily manage the growing visu…
cs.CV2026
Once-For-All: A Train-Once and Select-Anytime Framework for Multimodal Instruction Tuning
Mingkang Dong, Hongyi Cai, Xiwen Lei +3
Multimodal instruction tuning is the de facto recipe for adapting vision language models (VLMs), yet instruction data are highly redundant, making data selection critical for train…
cs.CV2025
AutoDebias: Automated Framework for Debiasing Text-to-Image Models
Hongyi Cai, Mohammad Mahdinur Rahman, Mingkang Dong +7
Text-to-Image (T2I) models generate high-quality images but are vulnerable to malicious backdoor attacks that inject harmful biases (e.g., trigger-activated gender or racial stereo…