3 papers
cs.CV2026
Natural Language Camera Movement Understanding
Yuwen Tan, Joey Huang, Jin Huang +2
Understanding camera movement in natural language is critical for training and evaluating video generation models, among other applications. However, we demonstrate that existing v…
cs.LG2025
LiteVLM: A Low-Latency Vision-Language Model Inference Pipeline for Resource-Constrained Environments
Jin Huang, Yuchao Jin, Le An +1
This paper introduces an efficient Vision-Language Model (VLM) pipeline specifically optimized for deployment on embedded devices, such as those used in robotics and autonomous dri…
cs.CV2025
VRsketch2Gaussian: 3D VR Sketch Guided 3D Object Generation with Gaussian Splatting
Songen Gu, Haoxuan Song, Binjie Liu +5
We propose VRSketch2Gaussian, a first VR sketch-guided, multi-modal, native 3D object generation framework that incorporates a 3D Gaussian Splatting representation. As part of our…