Publications (50)
SonoWorld: From One Image to a 3D Audio-Visual Scene
Derong Jin, Xiyi Chen, Ming C. Lin +1
AFFORD2ACT: Affordance-Guided Automatic Keypoint Selection for Generalizable and Lightweight Robotic Manipulation
Anukriti Singh, Kasra Torshizi, Khuzema Habib +3
Hearing Anywhere in Any Environment
Xiulong Liu, Anurag Kumar, Paul Calamia +7
Initial performance results of the JUNO detector
Angel Abusleme, Thomas Adam, Kai Adamowicz +1131
Geometry-Aware Multi-Task Learning for Binaural Audio Generation from Video
Rishabh Garg, Ruohan Gao, Kristen Grauman
ObjectFolder 2.0: A Multisensory Object Dataset for Sim2Real Transfer
Ruohan Gao, Zilin Si, Yen-Yu Chang +5
On-Demand Learning for Deep Image Restoration
Ruohan Gao, Kristen Grauman
Listen to Look: Action Recognition by Previewing Audio
Ruohan Gao, Tae-Hyun Oh, Kristen Grauman +1
Object-Centric Representation Learning from Unlabeled Videos
Ruohan Gao, Dinesh Jayaraman, Kristen Grauman
The ObjectFolder Benchmark: Multisensory Learning with Neural and Real Objects
Ruohan Gao, Yiming Dou, Hao Li +5
Learning to Highlight Audio by Watching Movies
Chao Huang, Ruohan Gao, J. M. F. Tsang +5
An Extensible Multimodal Multi-task Object Dataset with Materials
Trevor Standley, Ruohan Gao, Dawn Chen +2
Differentiable Physics Simulation of Dynamics-Augmented Neural Objects
Simon Le Cleac'h, Hong-Xing Yu, Michelle Guo +5
Do Audio-Visual Large Language Models Really See and Hear?
Ramaneswaran Selvakumar, Kaousheik Jayakumar, S Sakshi +3
2.5D Visual Sound
Ruohan Gao, Kristen Grauman
JULOC: A Local 3-D Refined Crust Model for the Geoneutrino Measurement at JUNO
Ruohan Gao, Zhiwei Li, Ran Han +7
ControlTac: Force- and Position-Controlled Tactile Data Augmentation with a Single Reference Image
Dongyu Luo, Kelin Yu, Amir-Hossein Shahidzadeh +3
VisualEchoes: Spatial Image Representation Learning through Echolocation
Ruohan Gao, Changan Chen, Ziad Al-Halah +2
SoundCam: A Dataset for Finding Humans Using Room Acoustics
Mason Wang, Samuel Clarke, Jui-Hsien Wang +2
Learning Object-Centric Neural Scattering Functions for Free-Viewpoint Relighting and Scene Composition
Hong-Xing Yu, Michelle Guo, Alireza Fathi +5
Sonicverse: A Multisensory Simulation Platform for Embodied Household Agents that See and Hear
Ruohan Gao, Hao Li, Gokul Dharan +6
Im2Flow: Motion Hallucination from Static Images for Action Recognition
Ruohan Gao, Bo Xiong, Kristen Grauman
Scene-wide Acoustic Parameter Estimation
Ricardo Falcon-Perez, Ruohan Gao, Gregor Mueckl +2
First measurement of reactor neutrino oscillations at JUNO
Angel Abusleme, Thomas Adam, Kai Adamowicz +1131
Aurelia: Test-time Reasoning Distillation in Audio-Visual LLMs
Sanjoy Chowdhury, Hanan Gani, Nishit Anand +5
NOIR: Neural Signal Operated Intelligent Robots for Everyday Activities
Ruohan Zhang, Sharon Lee, Minjune Hwang +11
Visual Acoustic Matching
Changan Chen, Ruohan Gao, Paul Calamia +1
ObjectFolder: A Dataset of Objects with Implicit Visual, Auditory, and Tactile Representations
Ruohan Gao, Yen-Yu Chang, Shivani Mall +2
VisualVoice: Audio-Visual Speech Separation with Cross-Modal Consistency
Ruohan Gao, Kristen Grauman
GenFlowRL: Shaping Rewards with Generative Object-Centric Flow in Visual Reinforcement Learning
Kelin Yu, Sheng Zhang, Harshit Soora +4
EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception
Sanjoy Chowdhury, Subrata Biswas, Sayan Nag +7
Hearing Anything Anywhere
Mason Wang, Ryosuke Sawata, Samuel Clarke +3
FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
Gabriele Oliaro, Xupeng Miao, Xinhao Cheng +9
OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies
Kelin Yu, Haode Zhang, Harish Ravichandar +2
DiffSound: Differentiable Modal Sound Rendering and Inverse Rendering for Diverse Inference Tasks
Xutong Jin, Chenxi Xu, Ruohan Gao +3
ShapeCodes: Self-Supervised Feature Learning by Lifting Views to Viewgrids
Dinesh Jayaraman, Ruohan Gao, Kristen Grauman
Co-Separating Sounds of Visual Objects
Ruohan Gao, Kristen Grauman
Learning to Set Waypoints for Audio-Visual Navigation
Changan Chen, Sagnik Majumder, Ziad Al-Halah +3
Prospects for geoneutrino detection with JUNO
Thomas Adam, Shakeel Ahmad, Rizwan Ahmed +625
Spherical World-Locking for Audio-Visual Localization in Egocentric Videos
Heeseung Yun, Ruohan Gao, Ishwarya Ananthabhotla +6
Towards Perception-Informed Latent HRTF Representations
You Zhang, Andrew Francl, Ruohan Gao +3
HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos
Zhi Wang, Botao He, Kelin Yu +4
Expected geoneutrino signal at JUNO using local integrated 3-D refined crustal model
Ran Han, ZhiWei Li, Ruohan Gao +11
Meerkat: Audio-Visual Large Language Model for Grounding in Space and Time
Sanjoy Chowdhury, Sayan Nag, Subhrajyoti Dasgupta +4
Differentiable Room Acoustic Rendering with Multi-View Vision Priors
Derong Jin, Ruohan Gao
The Audio-Visual Conversational Graph: From an Egocentric-Exocentric Perspective
Wenqi Jia, Miao Liu, Hao Jiang +4
See, Hear, and Feel: Smart Sensory Fusion for Robotic Manipulation
Hao Li, Yizhi Zhang, Junzhe Zhu +7
Learning to Separate Object Sounds by Watching Unlabeled Video
Ruohan Gao, Rogerio Feris, Kristen Grauman
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs
Sanjoy Chowdhury, Sayan Nag, Subhrajyoti Dasgupta +4
RealImpact: A Dataset of Impact Sound Fields for Real Objects
Samuel Clarke, Ruohan Gao, Mason Wang +5