collaborators

14 papers

cs.CV2026

DramaDirector: Geometry-Guided Short Drama Generation

Hengji Zhou, Sijie Liu, Jianrun Chen +3

Short dramas, with their rapid shot rhythms, dialogue-driven focus shifts, and demanding cinematographic grounding, pose challenges that prompt-level or text-only video generation…

cs.AI2026

Navigating User Behavior toward Personalized Multimodal Generation

Hengji Zhou, Yufeng Liu, Ye Liu +3

Modern AIGC pipelines deliver high-fidelity images and videos but presuppose a well-formed creation instruction, while end users rarely articulate visual details, leaving generator…

cs.AI2026

TailorMind: Towards Preference-Aligned Multimodal Content Generation

Hengji Zhou, Ye Liu, Yufeng Liu +3

Personalized content systems depend on available UGC and struggle when suitable content is absent, delayed, or costly to create. Although multimodal generators can synthesize conte…

cs.CV2026

TEMA: Anchor the Image, Follow the Text for Multi-Modification Composed Image Retrieval

Zixu Li, Yupeng Hu, Zhiheng Fu +3

Composed Image Retrieval (CIR) is an important image retrieval paradigm that enables users to retrieve a target image using a multimodal query that consists of a reference image an…

cs.CL2026

Parallel Test-Time Scaling for Latent Reasoning Models

Runyang You, Yongqi Li, Meng Liu +3

Parallel test-time scaling (TTS) is a pivotal approach for enhancing large language models (LLMs), typically by sampling multiple token-based chains-of-thought in parallel and aggr…

cs.AI2026

Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning

Dongjie Cheng, Yongqi Li, Zhixin Ma +5

Multimodal Large Language Models (MLLMs) are making significant progress in multimodal reasoning. Early approaches focus on pure text-based reasoning. More recent studies have inco…