From the 1 of 5 linked papers with an AI index.
5 papers
CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models
He Liang, Chenyang Ma, Yiming Zhang +4
The paper introduces CAIRN, a topology‑aware large multimodal model that uses graph neural networks and hierarchical attention to understand and reason about multi‑room 3D scenes.
A non-invasive video-based method for individual identification of wildlife using gait dynamics
Muhammad Aamir, Matthew Wijers, Sangyun Shin +2
Gait is a distinctive behavioral characteristic that enables non-invasive individual identification without requiring physical interaction with an animal. While gait-based analysis…
WildDepth: A Multimodal Dataset for 3D Wildlife Perception and Depth Estimation
Muhammad Aamir, Naoya Muramatsu, Sangyun Shin +6
Depth estimation and 3D reconstruction have been extensively studied as core topics in computer vision. Starting from rigid objects with relatively simple geometric shapes, such as…
DynPoint: Dynamic Neural Point For View Synthesis
Kaichen Zhou, Jia-Xing Zhong, Sangyun Shin +4
The introduction of neural radiance fields has greatly improved the effectiveness of view synthesis for monocular videos. However, existing algorithms face difficulties when dealin…
SoundLoc3D: Invisible 3D Sound Source Localization and Classification Using a Multimodal RGB-D Acoustic Camera
Yuhang He, Sangyun Shin, Anoop Cherian +2
Accurately localizing 3D sound sources and estimating their semantic labels -- where the sources may not be visible, but are assumed to lie on the physical surface of objects in th…