8 citations · 9 across the 13 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
VisGym: Diverse, Customizable, Scalable Environments for Multimodal Agents
Zirui Wang, Junyi Zhang, Jiaxin Ge +9
Modern Vision-Language Models (VLMs) remain poorly characterized in multi-step visual interactions, particularly in how they integrate perception, memory, and action over long hori…
cs.CV2024★ 8 cited
A Touch, Vision, and Language Dataset for Multimodal Alignment
Letian Fu, Gaurav Datta, Huang Huang +7
Touch is an important sensing modality for humans, but it has not yet been incorporated into a multimodal generative language model. This is partially due to the difficulty of obta…
cs.CV2024
Rethinking Patch Dependence for Masked Autoencoders
Letian Fu, Long Lian, Renhao Wang +6
In this work, we examine the impact of inter-patch dependencies in the decoder of masked autoencoders (MAE) on representation learning. We decompose the decoding mechanism for mask…