1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.CV2024★ 1 cited
Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos
Jianrui Zhang, Mu Cai, Yong Jae Lee
There has been growing sentiment recently that modern large multimodal models (LMMs) have addressed most of the key challenges related to short video comprehension. As a result, bo…
cs.CV2024
Cross-Modal Self-Supervised Learning with Effective Contrastive Units for LiDAR Point Clouds
Mu Cai, Chenxu Luo, Yong Jae Lee +1
3D perception in LiDAR point clouds is crucial for a self-driving vehicle to properly act in 3D environment. However, manually labeling point clouds is hard and costly. There has b…
cs.CV2024
VGBench: Evaluating Large Language Models on Vector Graphics Understanding and Generation
Bocheng Zou, Mu Cai, Jianrui Zhang +1
In the realm of vision models, the primary mode of representation is using pixels to rasterize the visual world. Yet this is not always the best or unique way to represent visual c…