5 papers
CoLoGen: Progressive Learning of Concept-Localization Duality for Unified Image Generation
YuXin Song, Yu Lu, Haoyuan Sun +6
Unified conditional image generation remains difficult because different tasks depend on fundamentally different internal representations. Some require conceptual understanding for…
Guiding Visual Autoregressive Models through Spectrum Weakening
Chaoyang Wang, Tianmeng Yang, Jingdong Wang +1
Classifier-free guidance (CFG) has become a widely adopted and practical approach for enhancing generation quality and improving condition alignment. Recent studies have explored g…
Perception Before Reasoning: Two-Stage Reinforcement Learning for Visual Reasoning in Vision-Language Models
Yan Chen, Long Li, Teng Xi +2
Reinforcement learning (RL) has proven highly effective in eliciting the reasoning capabilities of large language models (LLMs). Inspired by this success, recent studies have explo…
TEST-V: TEst-time Support-set Tuning for Zero-shot Video Classification
Rui Yan, Jin Wang, Hongyu Qu +4
Recently, adapting Vision Language Models (VLMs) to zero-shot visual classification by tuning class embedding with a few prompts (Test-time Prompt Tuning, TPT) or replacing class n…
Dense Connector for MLLMs
Huanjin Yao, Wenhao Wu, Taojiannan Yang +7
Do we fully leverage the potential of visual encoder in Multimodal Large Language Models (MLLMs)? The recent outstanding performance of MLLMs in multimodal understanding has garner…