Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
SceneBench: A Hierarchical Benchmark for Vision-Language Understanding of 3D Scenes
Anubhav Khanal, Prabigya Acharya, Roshni Poudel +5
Vision-language models excel at 2D image understanding but remain limited in 3D spatial reasoning. Progress is hindered by limitations in current benchmarks. First, 3D datasets oft…
cs.CV2024
iHuman: Instant Animatable Digital Humans From Monocular Videos
Pramish Paudel, Anubhav Khanal, Ajad Chhatkuli +2
Personalized 3D avatars require an animatable representation of digital humans. Doing so instantly from monocular videos offers scalability to broad class of users and wide-scale a…