From the 1 of 26 linked papers with an AI index.
8 papers · 1 filter
SceneBind: Binding What and Where Across Vision, Audio and Language
Mingfei Chen, Zijun Cui, Ruoke Zhang +2
SceneBind introduces an omni‑modal representation that jointly encodes what objects are and where they are in 3D space across vision, audio, and language, enabling cross‑modal scen…
Flowing from Reasoning to Motion: Learning 3D Hand Trajectory Prediction from Egocentric Human Interaction Videos
Mingfei Chen, Yifan Wang, Zhengqin Li +10
Prior works on 3D hand trajectory prediction are constrained by datasets that decouple motion from semantic supervision and by models that weakly link reasoning and action. To addr…
Neural Tangent Knowledge Distillation for Optical Convolutional Networks
Jinlin Xiang, Minho Choi, Yubo Zhang +3
Hybrid Optical Neural Networks (ONNs, typically consisting of an optical frontend and a digital backend) offer an energy-efficient alternative to fully digital deep networks for re…
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing
Mingfei Chen, Zijun Cui, Xiulong Liu +4
3D spatial reasoning in dynamic, audio-visual environments is a cornerstone of human cognition yet remains largely unexplored by existing Audio-Visual Large Language Models (AV-LLM…
Hearing Anywhere in Any Environment
Xiulong Liu, Anurag Kumar, Paul Calamia +7
In mixed reality applications, a realistic acoustic experience in spatial environments is as crucial as the visual experience for achieving true immersion. Despite recent advances…
Tell What You Hear From What You See -- Video to Audio Generation Through Text
Xiulong Liu, Kun Su, Eli Shlizerman
The content of visual and audio scenes is multi-faceted such that a video can be paired with various audio and vice-versa. Thereby, in video-to-audio generation task, it is imperat…