12 citations · 26 across the 52 of their papers we have counts for
53 papers
Decoupled Self-Forcing Distillation for Streaming Talking Head Generation
Yanru An, Ruiyan Wang, Wenwu Wei +7
Streaming talking-head generation produces each frame as its driving audio arrives, yet fidelity and efficiency have so far pulled in opposite directions: end-to-end methods condit…
EgoTac: In-the-wild Tactile Prediction from Egocentric Vision
Wenkang Zhang, Chengbo Yuan, Zicheng Zhang +2
Touch is fundamental to dexterous manipulation, yet most egocentric human data increasingly used for robot learning lacks tactile information. Directly collecting large-scale tacti…
AdvDex: Learning Dexterous Manipulation from Human Demonstrations via Joint-Aligned Actions and Adversarial Learning
Zhiyue Zhao, Jingyi Wu, Hairuo Liu +5
Dexterous manipulation is a fundamental capability for embodied intelligence, but scaling it remains difficult because robot demonstrations are expensive to collect and action spac…
Visual Token Codec: Unleashing Spatial Redundancy for ViT Feature Coding
Donghui Feng, Fengxi Zhang, Changsheng Gao +6
Distributed deployment of large vision foundation models often partitions a ViT backbone and exchanges intermediate token features between computing nodes, making efficient feature…
MIEScore: Human-Aligned Evaluation for Multi-Source Image Editing
Zitong Xu, Huiyu Duan, Xinyun Zhang +7
Recent advances in unified multimodal models have significantly improved text-guided image editing abilities. In particular, models such as Nano-Banana-Pro and GPT-Image-2 demonstr…
LLMCodec: Adapting Video Codecs for Efficient Weight Compression of Large Language Models
Rui Wang, Yan Zhao, Li Song +1
The rapid development of large language models(LLMs) has led to remarkable advances in natural language processing. However, the increasing scale of these models introduces substan…