Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
LLaVA-UHD v3: Progressive Visual Compression for Efficient Native-Resolution Encoding in MLLMs
Shichu Sun, Yichen Zhang, Haolin Song +6
Visual encoding followed by token condensing has become the standard architectural paradigm in multi-modal large language models (MLLMs). Many recent MLLMs increasingly favor globa…
cs.CV2025
Boosting Single-domain Generalized Object Detection via Vision-Language Knowledge Interaction
Xiaoran Xu, Jiangang Yang, Wenyue Chong +4
Single-Domain Generalized Object Detection~(S-DGOD) aims to train an object detector on a single source domain while generalizing well to diverse unseen target domains, making it s…
cs.CV2024
Object Gaussian for Monocular 6D Pose Estimation from Sparse Views
Luqing Luo, Shichu Sun, Jiangang Yang +3
Monocular object pose estimation, as a pivotal task in computer vision and robotics, heavily depends on accurate 2D-3D correspondences, which often demand costly CAD models that ma…