19 citations · 28 across the 5 of their papers we have counts for
5 papers
MobiFuse: A High-Precision On-device Depth Perception System with Multi-Data Fusion
Jinrui Zhang, Deyu Zhang, Tingting Long +6
We present MobiFuse, a high-precision depth perception system on mobile devices that combines dual RGB and Time-of-Flight (ToF) cameras. To achieve this, we leverage physical princ…
Ego3DPose: Capturing 3D Cues from Binocular Egocentric Views
Taeho Kang, Kyungjin Lee, Jinrui Zhang +1
We present Ego3DPose, a highly accurate binocular egocentric 3D pose reconstruction system. The binocular egocentric setup offers practicality and usefulness in various application…
Transferable Decoding with Visual Entities for Zero-Shot Image Captioning
Junjie Fei, Teng Wang, Jinrui Zhang +3
Image-to-text generation aims to describe images using natural language. Recently, zero-shot image captioning based on pre-trained vision-language models (VLMs) and large language…
Learning Grounded Vision-Language Representation for Versatile Understanding in Untrimmed Videos
Teng Wang, Jinrui Zhang, Feng Zheng +3
Joint video-language learning has received increasing attention in recent years. However, existing works mainly focus on single or multiple trimmed video clips (events), which make…
Exploiting Context Information for Generic Event Boundary Captioning
Jinrui Zhang, Teng Wang, Feng Zheng +2
Generic Event Boundary Captioning (GEBC) aims to generate three sentences describing the status change for a given time boundary. Previous methods only process the information of a…