collaborators

6 papers

cs.AI2026

Navigating User Behavior toward Personalized Multimodal Generation

Hengji Zhou, Yufeng Liu, Ye Liu +3

Modern AIGC pipelines deliver high-fidelity images and videos but presuppose a well-formed creation instruction, while end users rarely articulate visual details, leaving generator…

cs.AI2026

TailorMind: Towards Preference-Aligned Multimodal Content Generation

Hengji Zhou, Ye Liu, Yufeng Liu +3

Personalized content systems depend on available UGC and struggle when suitable content is absent, delayed, or costly to create. Although multimodal generators can synthesize conte…

cs.CV2026

VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning

Ye Liu, Kevin Qinghong Lin, Chang Wen Chen +1

Videos, with their unique temporal dimension, demand precise grounded understanding, where answers are directly linked to visual, interpretable evidence. Despite significant breakt…

cs.CV2025

UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning

Ye Liu, Zongyang Ma, Junfu Pu +4

Recent advances in Large Multi-modal Models (LMMs) have demonstrated their remarkable success as general-purpose multi-modal assistants, with particular focuses on holistic image-…

cs.CV2025

A Survey on Video Temporal Grounding with Multimodal Large Language Model

Jianlong Wu, Wei Liu, Ye Liu +4

The recent advancement in video temporal grounding (VTG) has significantly enhanced fine-grained video understanding, primarily driven by multimodal large language models (MLLMs).…

cs.CV2025

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos

Fanheng Kong, Jingyuan Zhang, Hongzhi Zhang +7

Videos are unique in their integration of temporal elements, including camera, scene, action, and attribute, along with their dynamic relationships over time. However, existing ben…