23 papers
EXPLORE-Bench: Egocentric Scene Prediction with Long-Horizon Reasoning
Chengjun Yu, Xuhan Zhu, Chaoqun Du +4
Multimodal large language models (MLLMs) are increasingly considered as a foundation for embodied agents, yet it remains unclear whether they can reliably reason about the long-ter…
Determinism of Randomness: Prompt-Residual Seed Shaping for Diffusion Generation
Song Yan, Wei Zhai, Chenfeng Wang +9
Diffusion models start generation from an isotropic Gaussian latent, yet changing only the random seed can lead to large differences in prompt faithfulness, composition, and visual…
End-to-End Spatial-Temporal Transformer for Real-time 4D HOI Reconstruction
Haoyu Zhang, Wei Zhai, Yuhang Yang +2
Monocular 4D human-object interaction (HOI) reconstruction - recovering a moving human and a manipulated object from a single RGB video - remains challenging due to depth ambiguity…
Event-based Visual Deformation Measurement
Yuliang Wu, Wei Zhai, Yuxin Cui +3
Visual Deformation Measurement (VDM) aims to recover dense deformation fields by tracking surface motion from camera observations. Traditional image-based methods rely on minimal i…
Towards Sequence Modeling Alignment between Tokenizer and Autoregressive Model
Pingyu Wu, Kai Zhu, Yu Liu +6
Autoregressive image generation aims to predict the next token based on previous ones. However, this process is challenged by the bidirectional dependencies inherent in conventiona…
Unbiased Gradient Estimation for Event Binning via Functional Backpropagation
Jinze Chen, Wei Zhai, Han Han +4
Event-based vision encodes dynamic scenes as asynchronous spatio-temporal spikes called events. To leverage conventional image processing pipelines, events are typically binned int…