17 papers
Balancing Efficiency and Efficacy: Training-Free Attention-Guided Switching Between Explicit and Latent Thoughts for MLLMs
Haoqian Kang, Liupeng Li, Kuofeng Gao +5
Reasoning in Multimodal Large Language Models (MLLMs) requires both fine-grained visual perception and rigorous logical deduction. Explicit text-based Chain-of-Thought (CoT) is com…
FlowErase-RL: Rethinking Concept Erasure as Reward Optimization in Flow Matching Models
Yi Sun, Zhiqi Zhang, Xinhao Zhong +5
Recent advances in flow matching models have significantly improved text-to-image generation quality, but also introduce growing safety risks due to the generation of harmful or un…
CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception
Liupeng Li, Haoqian Kang, Zhenyu Lu +4
High-resolution (HR) image perception presents a key bottleneck for multimodal large language models (MLLMs). While visual search offers a promising solution, existing methods stru…
CPC-VAR:Continual Personalized and Compositional Generation in Visual Autoregressive Models
Junhao Li, Xinhao Zhong, Yi sun +4
Visual autoregressive (VAR) models have recently emerged as an efficient paradigm for text-to-image generation. Despite their strong generative capability, existing VAR-based perso…
Prompt2Fingerprint: Plug-and-Play LLM Fingerprinting via Text-to-Weight Generation
Sixu Chen, Xiang Chen, Hongyao Yu +5
The widespread deployment and redistribution of large language models (LLMs) have made model provenance tracking a critical challenge. While existing LLM fingerprinting methods, pa…
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents
Niu Lian, Yuting Wang, Hanshu Yao +5
While multimodal large language models have demonstrated impressive short-term reasoning, they struggle with long-horizon video understanding due to limited context windows and sta…