3 papers
cs.CV2026
FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs
Eshika Khandelwal, Jingjing Pan, Mingfang Zhang +3
Multimodal large language models (MLLMs) are predominantly evaluated on free-form vision-language tasks such as visual question answering, captioning, and summarization. However, t…
cs.CV2026
SnapPose3D: Diffusion-Based Single-Frame 2D-to-3D Lifting of Human Poses
Alessandro Simoni, Riccardo Catalini, Davide Di Nucci +6
Depth ambiguity and joint uncertainty are the two main obstacles in obtaining accurate human pose predictions by 2D-to-3D lifting methods proposed in the literature. In particular,…
cs.CV2026
Object Pose Transformer: Unifying Unseen Object Pose Estimation
Weihang Li, Lorenzo Garattoni, Fabien Despinoy +2
Learning model-free object pose estimation for unseen instances remains a fundamental challenge in 3D vision. Existing methods typically fall into two disjoint paradigms: category-…