Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
EmbodiedHead: Real-Time Listening and Speaking Avatar for Conversational Agents
Yu Zhang, Kaiyuan Shen, Yang Li
We present EmbodiedHead, a speech-driven talking-head framework that equips LLMs with real-time visual avatars for conversation. A practical embodied avatar must achieve real-time…
cs.CV2026
HYDRA: Unifying Multi-modal Generation and Understanding via Representation-Harmonized Tokenization
Xuerui Qiu, Yutao Cui, Guozhen Zhang +9
Unified Multimodal Models struggle to bridge the fundamental gap between the abstract representations needed for visual understanding and the detailed primitives required for gener…
cs.CV2024
ChatTracker: Enhancing Visual Tracking Performance via Chatting with Multimodal Large Language Model
Yiming Sun, Fan Yu, Shaoxiang Chen +5
Visual object tracking aims to locate a targeted object in a video sequence based on an initial bounding box. Recently, Vision-Language~(VL) trackers have proposed to utilize addit…