collaborators

5 papers

cs.RO2026

OmniVTLA: Vision-Tactile-Language-Action Models with Semantic-Aligned Tactile Sensing

Zhengxue Cheng, Yiqian Zhang, Anni Tang +5

Recent vision-language-action (VLA) models build upon vision-language foundations, and have achieved promising results and exhibit the possibility of task generalization in robot m…

cs.CV2024

VidTok: A Versatile and Open-Source Video Tokenizer

Anni Tang, Tianyu He, Junliang Guo +3

Encoding video content into compact latent tokens has become a fundamental step in video generation and understanding, driven by the need to address the inherent redundancy in pixe…

cs.CV2024

Memories are One-to-Many Mapping Alleviators in Talking Face Generation

Anni Tang, Tianyu He, Xu Tan +2

Talking face generation aims at generating photo-realistic video portraits of a target person driven by input audio. Due to its nature of one-to-many mapping from the input audio t…

cs.MM2024

Rate-aware Compression for NeRF-based Volumetric Video

Zhiyu Zhang, Guo Lu, Huanxiong Liang +3

The neural radiance fields (NeRF) have advanced the development of 3D volumetric video technology, but the large data volumes they involve pose significant challenges for storage a…

cs.CV2024

Efficient Dynamic-NeRF Based Volumetric Video Coding with Rate Distortion Optimization

Zhiyu Zhang, Guo Lu, Huanxiong Liang +3

Volumetric videos, benefiting from immersive 3D realism and interactivity, hold vast potential for various applications, while the tremendous data volume poses significant challeng…