3 papers
cs.CV2026
KTV: Keyframes and Key Tokens Selection for Efficient Training-Free Video LLMs
Baiyang Song, Jun Peng, Yuxin Zhang +3
Training-free video understanding leverages the strong image comprehension capabilities of pre-trained vision language models (VLMs) by treating a video as a sequence of static fra…
cs.AI2025
Ming-Omni: A Unified Multimodal Model for Perception and Generation
Inclusion AI, Biao Gong, Cheng Zou +55
We propose Ming-Omni, a unified multimodal model capable of processing images, text, audio, and video, while demonstrating strong proficiency in both speech and image generation. M…
cs.CV2025
Multimodal Representation Learning Techniques for Comprehensive Facial State Analysis
Kaiwen Zheng, Xuri Ge, Junchen Fu +2
Multimodal foundation models have significantly improved feature representation by integrating information from multiple modalities, making them highly suitable for a broader set o…