2 papers
cs.CV2026
Progressive Multimodal Alignment for Continual Instruction Tuning
Duzhen Zhang, Yahan Yu, Qiaoyi Su +2
Multimodal Large Language Models (MLLMs) rely on a projector to align visual representations with the language embedding space, making it central to cross-modal understanding. In M…
cs.CV2025
Revisiting Continual Semantic Segmentation with Pre-trained Vision Models
Duzhen Zhang, Yong Ren, Wei Cong +9
Continual Semantic Segmentation (CSS) seeks to incrementally learn to segment novel classes while preserving knowledge of previously encountered ones. Recent advancements in CSS ha…