activity
20242026
collaborators

6 papers

cs.CV2026

SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM

An Yu, Weiheng Lu, Jian Li +4

Video Moment Retrieval is a task in video understanding that aims to localize a specific temporal segment in an untrimmed video based on a natural language query. Despite recent pr…

cs.LG2026

FlashSinkhorn: IO-Aware Entropic Optimal Transport on GPU

Felix X. -F. Ye, Xingjie Li, An Yu +3

Entropic optimal transport (EOT) via Sinkhorn iterations is widely used in modern machine learning, yet GPU solvers remain inefficient at scale. Tensorized implementations suffer q…

cs.CV2026

ReDiPrune: Relevance-Diversity Pre-Projection Token Pruning for Efficient Multimodal LLMs

An Yu, Ting Yu Tsai, Zhenfei Zhang +3

Recent multimodal large language models are computationally expensive because Transformers must process a large number of visual tokens. We present ReDiPrune, a training-free token…

eess.IV2025

IntelliCardiac: An Intelligent Platform for Cardiac Image Segmentation and Classification

Ting Yu Tsai, An Yu, Meghana Spurthi Maadugundu +5

Precise and effective processing of cardiac imaging data is critical for the identification and management of the cardiovascular diseases. We introduce IntelliCardiac, a comprehens…

cs.CV2025

LLM-Enabled Style and Content Regularization for Personalized Text-to-Image Generation

Anran Yu, Wei Feng, Yaochen Zhang +4

The personalized text-to-image generation has rapidly advanced with the emergence of Stable Diffusion. Existing methods, which typically fine-tune models using embedded identifiers…

cs.CV2024

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval

Weiheng Lu, Jian Li, An Yu +3

Multimodal Large Language Models (MLLMs) are widely used for visual perception, understanding, and reasoning. However, long video processing and precise moment retrieval remain cha…