3 papers
cs.CV2025
Perceptual Taxonomy: Evaluating and Guiding Hierarchical Scene Reasoning in Vision-Language Models
Jonathan Lee, Xingrui Wang, Jiawei Peng +9
We propose Perceptual Taxonomy, a structured process of scene understanding that first recognizes objects and their spatial configurations, then infers task-relevant properties suc…
cs.CV2025
4D-Animal: Freely Reconstructing Animatable 3D Animals from Videos
Shanshan Zhong, Jiawei Peng, Zehan Zheng +6
Existing methods for reconstructing animatable 3D animals from videos typically rely on sparse semantic keypoints to fit parametric models. However, obtaining such keypoints is lab…
cs.RO2025
PartInstruct: Part-level Instruction Following for Fine-grained Robot Manipulation
Yifan Yin, Zhengtao Han, Shivam Aarya +6
Fine-grained robot manipulation, such as lifting and rotating a bottle to display the label on the cap, requires robust reasoning about object parts and their relationships with in…