3 papers
cs.CV2026
TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens
Qingcheng Zhao, Yifang Pan, Karan Singh
Recent advances in Audio-LLMs like GPT-4o have ushered in an era of conversational interaction with language models. Conversational avatars however, still seem robotic in facial ex…
cs.CV2025
DepR: Depth Guided Single-view Scene Reconstruction with Instance-level Diffusion
Qingcheng Zhao, Xiang Zhang, Haiyang Xu +4
We propose DepR, a depth-guided single-view scene reconstruction framework that integrates instance-level diffusion within a compositional paradigm. Instead of reconstructing the e…
cs.CV2025
TANGLED: Generating 3D Hair Strands from Images with Arbitrary Styles and Viewpoints
Pengyu Long, Zijun Zhao, Min Ouyang +5
Hairstyles are intricate and culturally significant with various geometries, textures, and structures. Existing text or image-guided generation methods fail to handle the richness…