collaborators

6 papers

cs.CV2026

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models

Zeyu Xu, Xingzhong Hou, Pengkai Guo +6

Vision-Language Models (VLMs) have achieved strong progress in multimodal understanding. However, scaling dense or sparse Mixture-of-Experts (MoE) models to improve performance lim…

cs.CL2025

Online-PVLM: Advancing Personalized VLMs with Online Concept Learning

Huiyu Bai, Runze Wang, Zhuoyun Du +6

Personalized Visual Language Models (VLMs) are gaining increasing attention for their formidable ability in user-specific concepts aligned interactions (e.g., identifying a user's…

cs.CV2025

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning

Yi Liu, Xiao Xu, Zeyu Xu +10

Vision-Language Models (VLMs) have achieved remarkable breakthroughs in recent years, enabling a diverse array of applications in everyday life. However, the substantial computatio…

cs.CV2025

SWIFT: A General Sensitive Weight Identification Framework for Fast Sensor-Transfer Pansharpening

Zeyu Xia, Chenxi Sun, Tianyu Xin +3

Pansharpening aims to fuse high-resolution panchromatic (PAN) images with low-resolution multispectral (LRMS) images to generate high-resolution multispectral (HRMS) images. Althou…

cs.SD2025

Face2VoiceSync: Lightweight Face-Voice Consistency for Text-Driven Talking Face Generation

Fang Kang, Yin Cao, Haoyu Chen

Recent studies in speech-driven talking face generation achieve promising results, but their reliance on fixed-driven speech limits further applications (e.g., face-voice mismatch)…

cs.CL2025

Sounding Like a Winner? Prosodic Differences in Post-Match Interviews

Sofoklis Kakouros, Haoyu Chen

This study examines the prosodic characteristics associated with winning and losing in post-match tennis interviews. Additionally, this research explores the potential to classify…