most citedOpen-H-Embodiment: A Large-Scale Dataset for Enabling Foundation Models in Medical Robotics

1 citations · 1 across the 40 of their papers we have counts for

collaborators
Showing cs.CVShow all

29 papers · 1 filter

cs.CV2026

SOS! : A Streamlined Object-Conditional Transformer for Model-free Segmentation

Jiaqi Hu, Junwen Huang, Hongli Xu +4

Foundation segmentation models excel at generating high-quality, class-agnostic masks, but they struggle to associate these proposals with specific target objects. This semantic ga…

cs.CV2026

OSCAR: Occupancy-based Shape Completion via Acoustic Neural Implicit Representations

Magdalena Wysocki, Kadir Burak Buldu, Miruna-Alexandra Gafencu +2

Accurate 3D reconstruction of vertebral anatomy from ultrasound is important for guiding minimally invasive spine interventions, but it remains challenging due to acoustic shadowin…

cs.CV2026

HyperVLP: Enhancing Hierarchical Surgical Video-Language Pre-training in Hyperbolic Space

Yaojun Hu, Kun Yuan, Nassir Navab +3

Surgical vision-language foundation models typically adopt educational materials, such as surgical lecture videos, to transfer surgical knowledge encoded in language into visual re…

cs.CV2026

Pose Anything Anywhere:Model-free Object Poses from Arbitrary References

Hongli Xu, Jiaqi Hu, Junwen Huang +5

Estimating the 6D pose of unseen objects is a fundamental yet challenging problem for open-world robotics and embodied perception. Model-based methods are accurate but depend on CA…

cs.CV2026

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture

Felix Tristram, Stefano Gasperini, Benjamin Killeen +4

The increasing maturity of embodied AI platforms has driven a growing interest in procedural video representation learning to support intelligent assistance systems for complex, mu…

cs.CV2026

Prompting Diffusion Models for Zero-Shot Instance Segmentation

Irem Zeynep Alagöz, Nils Morbitzer, Andrea Ramazzina +3

Several disruptive research directions have recently emerged in computer vision, including foundation models achieving previously unseen zero-shot performance in scene understandin…