activity
20242026
most citedImproved Video VAE for Latent Video Diffusion Model

1 citations · 1 across the 11 of their papers we have counts for

collaborators

24 papers

cs.CV2026

End-to-End Spatial-Temporal Transformer for Real-time 4D HOI Reconstruction

Haoyu Zhang, Wei Zhai, Yuhang Yang +2

Monocular 4D human-object interaction (HOI) reconstruction - recovering a moving human and a manipulated object from a single RGB video - remains challenging due to depth ambiguity…

cs.CV2026

Event-based Visual Deformation Measurement

Yuliang Wu, Wei Zhai, Yuxin Cui +3

Visual Deformation Measurement (VDM) aims to recover dense deformation fields by tracking surface motion from camera observations. Traditional image-based methods rely on minimal i…

cs.CV2026

Unbiased Gradient Estimation for Event Binning via Functional Backpropagation

Jinze Chen, Wei Zhai, Han Han +4

Event-based vision encodes dynamic scenes as asynchronous spatio-temporal spikes called events. To leverage conventional image processing pipelines, events are typically binned int…

cs.CV2026

OmniVaT: Single Domain Generalization for Multimodal Visual-Tactile Learning

Liuxiang Qiu, Hui Da, Yuzhen Niu +3

Visual-tactile learning (VTL) enables embodied agents to perceive the physical world by integrating visual (VIS) and tactile (TAC) sensors. However, VTL still suffers from modality…

cs.LG2025

Anchoring Values in Temporal and Group Dimensions for Flow Matching Model Alignment

Yawen Shao, Jie Xiao, Kai Zhu +4

Group Relative Policy Optimization (GRPO) has proven highly effective in enhancing the alignment capabilities of Large Language Models (LLMs). However, current adaptations of GRPO…

cs.CV2025

TOUCH: Text-guided Controllable Generation of Free-Form Hand-Object Interactions

Guangyi Han, Wei Zhai, Yuhang Yang +2

Hand-object interaction (HOI) is fundamental for humans to express intent. Existing HOI generation research is predominantly confined to fixed grasping patterns, where control is t…