Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
VLGA: Vision-Language-Geometry-Action Models for Autonomous Driving
Jin Yao, Dhruva Dixith Kurra, Tom Lampo +3
Vision-language-action (VLA) models can describe scenes and reason about them in language, yet still struggle to ground their actions in the dense 3D world around them. Existing ap…
cs.CV2025
KineDiff3D: Kinematic-Aware Diffusion for Category-Level Articulated Object Shape Reconstruction and Generation
WenBo Xu, Liu Liu, Li Zhang +4
Articulated objects, such as laptops and drawers, exhibit significant challenges for 3D reconstruction and pose estimation due to their multi-part geometries and variable joint con…
cs.CV2025
Distilling Textual Priors from LLM to Efficient Image Fusion
Ran Zhang, Xuanhua He, Ke Cao +4
Multi-modality image fusion aims to synthesize a single, comprehensive image from multiple source inputs. Traditional approaches, such as CNNs and GANs, offer efficiency but strugg…