9 papers
SAB-LVLM: Significance-Aware Binarization for Large Vision-Language Models
Qi Lyu, Jiahua Dong, Baichen Liu +7
Large Vision-Language Models (LVLMs) have achieved remarkable progress in multimodal understanding, yet their enormous parameter scale and cross-modal computation incur substantial…
Does YOLO Really Need to See Every Training Image in Every Epoch?
Xingxing Xie, Jiahua Dong, Junwei Han +1
YOLO detectors are known for their fast inference speed, yet training them remains unexpectedly time-consuming due to their exhaustive pipeline that processes every training image…
Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
Pengfei Hu, Meng Cao, Yingyao Wang +6
Long video understanding is essential for human-like intelligence, enabling coherent perception and reasoning over extended temporal contexts. While the emerging thinking-with-fram…
Bring Your Dreams to Life: Continual Text-to-Video Customization
Jiahua Dong, Xudong Wang, Wenqi Liang +7
Customized text-to-video generation (CTVG) has recently witnessed great progress in generating tailored videos from user-specific text. However, most CTVG methods assume that perso…
CE-SDWV: Effective and Efficient Concept Erasure for Text-to-Image Diffusion Models via a Semantic-Driven Word Vocabulary
Jiahang Tu, Qian Feng, Jiahua Dong +4
Large-scale text-to-image (T2I) diffusion models have achieved remarkable generative performance about various concepts. With the limitation of privacy and safety in practice, the…
CRISP: Contrastive Residual Injection and Semantic Prompting for Continual Video Instance Segmentation
Baichen Liu, Qi Lyu, Xudong Wang +3
Continual video instance segmentation demands both the plasticity to absorb new object categories and the stability to retain previously learned ones, all while preserving temporal…