11 papers
Enhancing Video Physical Consistency via Role-aware Joint Training and Modality-decoupled Denoising
Guangting Zheng, Haojing Chen, Hao Li +6
While modern video diffusion models excel in visual fidelity, maintaining long-range physical consistency remains a formidable challenge. Conventional pixel-reconstruction objectiv…
Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation
Che Liu, Lichao Ma, Xiangyu Tony Zhang +4
Omni-modal language models are intended to jointly understand audio, visual inputs, and language, but benchmark gains can be inflated when visual evidence alone is enough to answer…
SafeSteer: A Decoding-level Defense Mechanism for Multimodal Large Language Models
Xinyi Zeng, Xue Yang, Jingyuan Zhang +5
Multimodal large language models (MLLMs) are gaining increasing attention. Due to the heterogeneity of their input features, they face significant challenges in terms of jailbreak…
SWIFT: Prompt-Adaptive Memory for Efficient Interactive Long Video Generation
Shanwen Tan, Hao Li, Jingtao Zhang +4
Streaming long-video generation faces a central challenge in continuous semantic switching, requiring adaptive memory to preserve coherent visual evolution. Current approaches rely…
DPN-LE: Dual Personality Neuron Localization and Editing for Large Language Models
Lifan Zheng, Xue Yang, Jiawei Chen +6
With the widespread adoption of large language models (LLMs), understanding their personality representation mechanisms has become critical. As a novel paradigm in Personality Edit…
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
Gen Luo, Wenhan Dou, Wenhao Li +9
This paper focuses on monolithic Multimodal Large Language Models (MLLMs), which integrate visual encoding and language decoding into a single model. Existing structures and pre-tr…