3 papers
cs.CV2026
FineViT: Progressively Unlocking Fine-Grained Perception with Dense Recaptions
Peisen Zhao, Xiaopeng Zhang, Mingxing Xu +10
While Multimodal Large Language Models (MLLMs) have experienced rapid advancements, their visual encoders frequently remain a performance bottleneck. Conventional CLIP-based encode…
cs.CV2024
Text-Animator: Controllable Visual Text Video Generation
Lin Liu, Quande Liu, Shengju Qian +5
Video generation is a challenging yet pivotal task in various industries, such as gaming, e-commerce, and advertising. One significant unresolved aspect within T2V is the effective…
cs.CV2024
Controllable Relation Disentanglement for Few-Shot Class-Incremental Learning
Yuan Zhou, Richang Hong, Yanrong Guo +3
In this paper, we propose to tackle Few-Shot Class-Incremental Learning (FSCIL) from a new perspective, i.e., relation disentanglement, which means enhancing FSCIL via disentanglin…