activity
20242026
collaborators

6 papers

cs.CV2026

ApET: Approximation-Error Guided Token Compression for Efficient VLMs

Qiankun Ma, Ziyao Zhang, Haofei Wang +3

Recent Vision-Language Models (VLMs) have demonstrated remarkable multimodal understanding capabilities, yet the redundant visual tokens incur prohibitive computational overhead an…

cs.CV2025

UniAPO: Unified Multimodal Automated Prompt Optimization

Qipeng Zhu, Yanzhe Chen, Huasong Zhong +5

Prompting is fundamental to unlocking the full potential of large language models. To automate and enhance this process, automatic prompt optimization (APO) has been developed, dem…

cs.CV2025

EF-VI: Enhancing End-Frame Injection for Video Inbetweening

Liuhan Chen, Xiaodong Cun, Xiaoyu Li +5

Video inbetweening aims to synthesize intermediate video sequences conditioned on the given start and end frames. Current state-of-the-art methods primarily extend large-scale pre-…

cs.CV2025

MV-CoRe: Multimodal Visual-Conceptual Reasoning for Complex Visual Question Answering

Jingwei Peng, Jiehao Chen, Mateo Alejandro Rojas +1

Complex Visual Question Answering (Complex VQA) tasks, which demand sophisticated multi-modal reasoning and external knowledge integration, present significant challenges for exist…

cs.GR2025

TRiMM: Transformer-Based Rich Motion Matching for Real-Time multi-modal Interaction in Digital Humans

Yueqian Guo, Tianzhao Li, Xin Lyu +7

Large Language Model (LLM)-driven digital humans have sparked a series of recent studies on co-speech gesture generation systems. However, existing approaches struggle with real-ti…

cs.CV2024

Knowing Where to Focus: Attention-Guided Alignment for Text-based Person Search

Lei Tan, Weihao Li, Pingyang Dai +3

In the realm of Text-Based Person Search (TBPS), mainstream methods aim to explore more efficient interaction frameworks between text descriptions and visual data. However, recent…