2 papers
cs.CV2025
DEMO: Disentangled Motion Latent Flow Matching for Fine-Grained Controllable Talking Portrait Synthesis
Peiyin Chen, Zhuowei Yang, Hui Feng +2
Audio-driven talking-head generation has advanced rapidly with diffusion-based generative models, yet producing temporally coherent videos with fine-grained motion control remains…
cs.CV2025
LVLMs as inspectors: an agentic framework for category-level structural defect annotation
Sheng Jiang, Yuanmin Ning, Bingxi Huang +2
Automated structural defect annotation is essential for ensuring infrastructure safety while minimizing the high costs and inefficiencies of manual labeling. A novel agentic annota…