49 papers
EgoTac: In-the-wild Tactile Prediction from Egocentric Vision
Wenkang Zhang, Chengbo Yuan, Zicheng Zhang +2
Touch is fundamental to dexterous manipulation, yet most egocentric human data increasingly used for robot learning lacks tactile information. Directly collecting large-scale tacti…
E-S2Feat:Semantic-Guided Spiking Local Feature Detection and Description for Event Cameras
Yang Yi, Juntao Hua, Jinpu Zhang +3
Benefiting from high temporal resolution and dynamic range, event-based local feature methods have attracted increasing attention. However, event sparsity, noise, and limited textu…
Visual Token Codec: Unleashing Spatial Redundancy for ViT Feature Coding
Donghui Feng, Fengxi Zhang, Changsheng Gao +6
Distributed deployment of large vision foundation models often partitions a ViT backbone and exchanges intermediate token features between computing nodes, making efficient feature…
MIEScore: Human-Aligned Evaluation for Multi-Source Image Editing
Zitong Xu, Huiyu Duan, Xinyun Zhang +7
Recent advances in unified multimodal models have significantly improved text-guided image editing abilities. In particular, models such as Nano-Banana-Pro and GPT-Image-2 demonstr…
LLMCodec: Adapting Video Codecs for Efficient Weight Compression of Large Language Models
Rui Wang, Yan Zhao, Li Song +1
The rapid development of large language models(LLMs) has led to remarkable advances in natural language processing. However, the increasing scale of these models introduces substan…
OmniVTLA: Vision-Tactile-Language-Action Models with Semantic-Aligned Tactile Sensing
Zhengxue Cheng, Yiqian Zhang, Anni Tang +5
Recent vision-language-action (VLA) models build upon vision-language foundations, and have achieved promising results and exhibit the possibility of task generalization in robot m…