3 papers
cs.CV2026
TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens
Qingcheng Zhao, Yifang Pan, Karan Singh
Recent advances in Audio-LLMs like GPT-4o have ushered in an era of conversational interaction with language models. Conversational avatars however, still seem robotic in facial ex…
cs.GR2026
Learning to Build Shapes by Extrusion
Thor Vestergaard Christiansen, Karran Pandey, Alba Reinders +3
We introduce Text Encoded Extrusions (TEE), a text-based representation that expresses mesh construction as sequences of face extrusions rather than polygon lists, and a method for…
cs.CV2024
Motion Modes: What Could Happen Next?
Karran Pandey, Matheus Gadelha, Yannick Hold-Geoffroy +3
Predicting diverse object motions from a single static image remains challenging, as current video generation models often entangle object movement with camera motion and other sce…