Publications (17)
Primitive Shape Recognition for Object Grasping
Yunzhi Lin, Chao Tang, Fu-Jen Chu +2
Shape informs how an object should be grasped, both in terms of where and how. As such, this paper describes a segmentation-based architecture for decomposing objects sensed with a…
Using Synthetic Data and Deep Networks to Recognize Primitive Shapes for Object Grasping
Yunzhi Lin, Chao Tang, Fu-Jen Chu +1
A segmentation-based architecture is proposed to decompose objects into multiple primitive shapes from monocular depth input for robotic manipulation. The backbone deep network is…
Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
Runsen Xu, Weiyao Wang, Hao Tang +5
Multi-modal large language models (MLLMs) have rapidly advanced in visual tasks, yet their spatial understanding remains limited to single images, leaving them ill-suited for physi…
Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
Kristen Grauman, Andrew Westbury, Lorenzo Torresani +98
We present Ego-Exo4D, a diverse, large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric…
HyperMix: Out-of-Distribution Detection and Classification in Few-Shot Settings
Nikhil Mehta, Kevin J Liang, Jing Huang +3
Out-of-distribution (OOD) detection is an important topic for real-world machine learning systems, but settings with limited in-distribution samples have been underexplored. Such f…
Supervised Collective Classification for Crowdsourcing
Pin-Yu Chen, Chia-Wei Lien, Fu-Jen Chu +2
Crowdsourcing utilizes the wisdom of crowds for collective classification via information (e.g., labels of an item) provided by labelers. Current crowdsourcing algorithms are mainl…