most citedTemplate-Free Single-View 3D Human Digitalization with Diffusion-Guided LRM

2 citations · 4 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CV2024

DeforHMR: Vision Transformer with Deformable Cross-Attention for 3D Human Mesh Recovery

Jaewoo Heo, George Hu, Zeyu Wang +1

Human Mesh Recovery (HMR) is an important yet challenging problem with applications across various domains including motion capture, augmented reality, and biomechanics. Accurately…

cs.CV2024

Motion Diffusion-Guided 3D Global HMR from a Dynamic Camera

Jaewoo Heo, Kuan-Chieh Wang, Karen Liu +1

Motion capture technologies have transformed numerous fields, from the film and gaming industries to sports science and healthcare, by providing a tool to capture and analyze human…

cs.CV2024

Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

Orr Zohar, Xiaohan Wang, Yonatan Bitton +2

The performance of Large Vision Language Models (LVLMs) is dependent on the size and quality of their training datasets. Existing video instruction tuning datasets lack diversity a…

cs.CV2024

μ-Bench: A Vision-Language Benchmark for Microscopy Understanding

Alejandro Lozano, Jeffrey Nirschl, James Burgess +4

Recent advances in microscopy have enabled the rapid generation of terabytes of image data in cell biology and biomedical research. Vision-language models (VLMs) offer a promising…

cs.CV20241 cited

VideoAgent: Long-form Video Understanding with Large Language Model as Agent

Xiaohan Wang, Yuhui Zhang, Orr Zohar +1

Long-form video understanding represents a significant challenge within computer vision, demanding a model capable of reasoning over long multi-modal sequences. Motivated by the hu…

cs.LG20241 cited

Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data

Yuhui Zhang, Elaine Sui, Serena Yeung-Levy

Building cross-modal applications is challenging due to limited paired multi-modal data. Recent works have shown that leveraging a pre-trained multi-modal contrastive representatio…