collaborators

6 papers

cs.CV2026

TimePLE: Rethinking Temporal Representation for Video Temporal Grounding

Yuhui Zeng, Xinyu Mao, Xiaokun Liu +4

Video temporal grounding (VTG) aims to localize the continuous video interval described by a natural-language query. However, current VLM-based methods typically produce this inter…

cs.AI2026

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning

Xinyu Mao, Yuhui Zeng, Xiaokun Liu +6

Cinematographic captioning aims to describe how a video is filmed using professional film-language concepts such as camera movement, shot size, depth of field, composition, and sho…

cs.CV2026

Fast-SAM3D: 3Dfy Anything in Images but Faster

Weilun Feng, Mingqiang Wu, Zhiliang Chen +10

SAM3D enables scalable, open-world 3D reconstruction from complex scenes, yet its deployment is hindered by prohibitive inference latency. In this work, we conduct the \textbf{firs…

cs.CV2025

FilmWeaver: Weaving Consistent Multi-Shot Videos with Cache-Guided Autoregressive Diffusion

Xiangyang Luo, Qingyu Li, Xiaokun Liu +7

Current video generation models perform well at single-shot synthesis but struggle with multi-shot videos, facing critical challenges in maintaining character and background consis…

cs.CV2025

Improving Video Generation with Human Feedback

Jie Liu, Gongye Liu, Jiajun Liang +14

Video generation has achieved significant advances through rectified flow techniques, but issues like unsmooth motion and misalignment between videos and prompts persist. In this w…

cs.CV2025

HumanAesExpert: Advancing a Multi-Modality Foundation Model for Human Image Aesthetic Assessment

Zhichao Liao, Xiaokun Liu, Wenyu Qin +6

Image Aesthetic Assessment (IAA) is a long-standing and challenging research task. However, its subset, Human Image Aesthetic Assessment (HIAA), has been scarcely explored. To brid…