12 papers
VideoVIBE: A Video-Grounded Diagnostic Benchmark for One-Shot Interactive Website Generation
Jiajun Xu, Yanghao Zhou, Jingyun Liao +6
Natural-language-driven "vibe coding" enables the one-shot generation of visually rich and interactive web applications, yet reliable assessment of their quality has not kept pace.…
MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation
Haitian Li, Yanghao Zhou, Heyan Huang +15
In recent years, Multi-Talker Audio-Video Generation (MTAVG) models have shown promising performance on fundamental metrics such as lip-sync and audio-visual alignment. However, th…
Exploring a Multimodal Chatbot as a Facilitator in Therapeutic Art Activity
Le Lin, Zihao Zhu, Rainbow Tin Hung Ho +2
Therapeutic art activities, such as expressive drawing and painting, require the synergy between creative visual production and interactive dialogue. Recent advancements in Multimo…
MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation
Yang-Hao Zhou, Haitian Li, Rexar Lin +12
Recent advances in text-to-audio-video (T2AV) generation have enabled models to synthesize audio-visual videos with multi-participant dialogues. However, existing evaluation benchm…
Distractor-free Generalizable 3D Gaussian Splatting
Yanqi Bao, Jing Liao, Jing Huo +1
We present DGGS, a novel framework that addresses the previously unexplored challenge: (3DGS). It mitigates 3D incons…
TalkingEyes: Pluralistic Speech-Driven 3D Eye Gaze Animation
Yixiang Zhuang, Chunshan Ma, Yao Cheng +3
Although significant progress has been made in the field of speech-driven 3D facial animation recently, the speech-driven animation of an indispensable facial component, eye gaze,…