collaborators

6 papers

cs.CV2026

LiveGesture Streamable Co-Speech Gesture Generation Model

Muhammad Usama Saleem, Mayur Jagdishbhai Patel, Ekkasit Pinyoanuntapong +6

We propose LiveGesture, the first fully streamable, speech-driven full-body gesture generation framework that operates with zero look-ahead and supports arbitrary sequence length.…

cs.CV2026

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization

Yuqi Chen, Xiaohan Zhang, Ahmad Arrabi +3

Natural-language Guided Cross-view Geo-localization (NGCG) aims to retrieve geo-tagged satellite imagery using textual descriptions of ground scenes. While recent NGCG methods comm…

cs.CV2026

How Should Video LLMs Output Time? An Analysis of Efficient Temporal Grounding Paradigms

Shengji Jin, Yuanhao Zou, Victor Zhu +2

While Multimodal Large Language Models (MLLMs) have advanced Video Temporal Grounding (VTG), existing methods often couple output paradigms with different backbones, datasets, and…

cs.CV2025

Autoregressive Video Generation beyond Next Frames Prediction

Sucheng Ren, Chen Chen, Zhenbang Wang +5

Autoregressive models for video generation typically operate frame-by-frame, extending next-token prediction from language to video's temporal dimension. We question that unlike wo…

cs.CV2025

MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer

Yanghao Li, Rui Qian, Bowen Pan +24

Unified multimodal Large Language Models (LLMs) that can both understand and generate visual content hold immense potential. However, existing open-source models often suffer from…

cs.CV2025

AToken: A Unified Tokenizer for Vision

Jiasen Lu, Liangchen Song, Mingze Xu +5

We present AToken, the first unified visual tokenizer that achieves both high-fidelity reconstruction and semantic understanding across images, videos, and 3D assets. Unlike existi…