activity
20242026
most citedXYZ-IBD: Benchmarking Robust 6D Object Pose Estimation under Real-World Industrial Complexity

2 citations · 2 across the 2 of their papers we have counts for

collaborators

20 papers

cs.CV2026

When does fusing hand-crafted knowledge with learned representations pay? A cost-normalized benchmark of stacking, substitution, and interference

Ahmad AlMughrabi, Albert Clop, Benjamin Busam +2

Fusing prior knowledge with data-driven learning is attractive where data is scarce, yet no controlled account says when it helps, is redundant, or harms. We benchmark one fixed ha…

cs.CV20262 cited

XYZ-IBD: Benchmarking Robust 6D Object Pose Estimation under Real-World Industrial Complexity

Junwen Huang, Jiaqi Hu, Peter KT Yu +3

While current 6D pose estimation benchmarks have reached near-saturation on household objects, they often fail to capture the stochastic and optical complexities of industrial envi…

cs.CV2026

Pose Anything Anywhere:Model-free Object Poses from Arbitrary References

Hongli Xu, Jiaqi Hu, Junwen Huang +5

Estimating the 6D pose of unseen objects is a fundamental yet challenging problem for open-world robotics and embodied perception. Model-based methods are accurate but depend on CA…

cs.GR2026

PromptVFX: Text-Driven Fields for Open-World 3D Gaussian Animation

Mert Kiray, Paul Uhlenbruck, Nassir Navab +1

Visual effects (VFX) are key to immersion in modern films, games, and AR/VR. Creating 3D effects requires specialized expertise and training in 3D animation software and can be tim…

cs.CV2026

BenchSeg: A Large-Scale Dataset and Benchmark for Multi-View Food Video Segmentation

Ahmad AlMughrabi, Guillermo Rivo, Carlos Jiménez-Farfán +6

Food image segmentation is a critical task for dietary analysis, enabling accurate estimation of food volume and nutrients. However, current methods suffer from limited multi-view…

cs.RO2025

STARE-VLA: Progressive Stage-Aware Reinforcement for Fine-Tuning Vision-Language-Action Models

Feng Xu, Guangyao Zhai, Xin Kong +4

Recent advances in Vision-Language-Action (VLA) models, powered by large language models and reinforcement learning-based fine-tuning, have shown remarkable progress in robotic man…