most citedPartFormer: Awakening Latent Diverse Representation from Vision Transformer for Object Re-Identification

1 citations · 2 across the 5 of their papers we have counts for

collaborators

6 papers

cs.CV2025

UniAPO: Unified Multimodal Automated Prompt Optimization

Qipeng Zhu, Yanzhe Chen, Huasong Zhong +5

Prompting is fundamental to unlocking the full potential of large language models. To automate and enhance this process, automatic prompt optimization (APO) has been developed, dem…

cs.CV2025

MV-CoRe: Multimodal Visual-Conceptual Reasoning for Complex Visual Question Answering

Jingwei Peng, Jiehao Chen, Mateo Alejandro Rojas +1

Complex Visual Question Answering (Complex VQA) tasks, which demand sophisticated multi-modal reasoning and external knowledge integration, present significant challenges for exist…

cs.GR20251 cited

TRiMM: Transformer-Based Rich Motion Matching for Real-Time multi-modal Interaction in Digital Humans

Yueqian Guo, Tianzhao Li, Xin Lyu +7

Large Language Model (LLM)-driven digital humans have sparked a series of recent studies on co-speech gesture generation systems. However, existing approaches struggle with real-ti…

cs.CV2025

EF-VI: Enhancing End-Frame Injection for Video Inbetweening

Liuhan Chen, Xiaodong Cun, Xiaoyu Li +5

Video inbetweening aims to synthesize intermediate video sequences conditioned on the given start and end frames. Current state-of-the-art methods primarily extend large-scale pre-…

cs.CV2024

Knowing Where to Focus: Attention-Guided Alignment for Text-based Person Search

Lei Tan, Weihao Li, Pingyang Dai +3

In the realm of Text-Based Person Search (TBPS), mainstream methods aim to explore more efficient interaction frameworks between text descriptions and visual data. However, recent…

cs.CV20241 cited

PartFormer: Awakening Latent Diverse Representation from Vision Transformer for Object Re-Identification

Lei Tan, Pingyang Dai, Jie Chen +3

Extracting robust feature representation is critical for object re-identification to accurately identify objects across non-overlapping cameras. Although having a strong representa…