3 papers
cs.CV2026
Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation
Dazhao Du, Shiyan Du, Jian Liu +8
Understanding camera motion is fundamental to video perception, with applications in spatial intelligence and controllable video generation. Multimodal large language models (MLLMs…
cs.CV2026
DEIG: Detail-Enhanced Instance Generation with Fine-Grained Semantic Control
Shiyan Du, Conghan Yue, Xinyu Cheng +1
Multi-Instance Generation has advanced significantly in spatial placement and attribute binding. However, existing approaches still face challenges in fine-grained semantic underst…
cs.CV2026
Improving Fine-Grained Control via Aggregation of Multiple Diffusion Models
Conghan Yue, Zhengwei Peng, Shiyan Du +4
While many diffusion models perform well when controlling particular aspects such as style, character, and interaction, they struggle with fine-grained control due to dataset limit…