7 papers
InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos
Chi Zhang, Haoyang Shi, Yueyi Liu +4
Large language models have made text the default medium for human--AI interaction, buttext alone cannot express the full range of responses required by multimodal assistants,avatar…
ElasticTTT: Prior-Preserving Test-Time Tuning for Video Editing
Yueyi Liu, Chi Zhang, Sen Cui +1
Test-Time Tuning (TTT) on pretrained diffusion models has emerged as a powerful paradigm for video editing. However, there exists a foundational mismatch between the distribution-m…
PIPBench: A Profile-Inclusive Framework for Personalized Image Generation Evaluation
Yuhang Wu, Shuxiang Zhang, Wee Hian Ching +2
Recent text-to-image models such as DALLE-3 excel at following diverse prompts yet remain blind to individual aesthetic preferences. We study personalized image generation, where m…
Disciplined Diffusion: Text-to-Image Diffusion Model against NSFW Generation
Chi Zhang, Changjia Zhu, Xiaowen Li +2
Text-to-image (T2I) diffusion models have the ability to build high-quality pictures from text prompts, but they pose safety concerns because they can generate offensive or disturb…
X2SAM: Any Segmentation in Images and Videos
Hao Wang, Limeng Qiao, Chi Zhang +4
Multimodal Large Language Models (MLLMs) have demonstrated strong image-level visual understanding and reasoning, yet their pixel-level perception across both images and videos rem…
StyleStudio: Text-Driven Style Transfer with Selective Control of Style Elements
Mingkun Lei, Xue Song, Beier Zhu +2
Text-driven style transfer aims to merge the style of a reference image with content described by a text prompt. Recent advancements in text-to-image models have improved the nuanc…