NewEvery arXiv paper, its researchers & institutions — mapped.
papers

Publications (50)

cs.CV2026

SonoWorld: From One Image to a 3D Audio-Visual Scene

Derong Jin, Xiyi Chen, Ming C. Lin +1

cs.RO2026

AFFORD2ACT: Affordance-Guided Automatic Keypoint Selection for Generalizable and Lightweight Robotic Manipulation

Anukriti Singh, Kasra Torshizi, Khuzema Habib +3

cs.CV2025

Hearing Anywhere in Any Environment

Xiulong Liu, Anurag Kumar, Paul Calamia +7

hep-ex2025

Initial performance results of the JUNO detector

Angel Abusleme, Thomas Adam, Kai Adamowicz +1131

cs.CV2021

Geometry-Aware Multi-Task Learning for Binaural Audio Generation from Video

Rishabh Garg, Ruohan Gao, Kristen Grauman

cs.CV2022

ObjectFolder 2.0: A Multisensory Object Dataset for Sim2Real Transfer

Ruohan Gao, Zilin Si, Yen-Yu Chang +5

cs.CV2017

On-Demand Learning for Deep Image Restoration

Ruohan Gao, Kristen Grauman

cs.CV2020

Listen to Look: Action Recognition by Previewing Audio

Ruohan Gao, Tae-Hyun Oh, Kristen Grauman +1

cs.CV2016

Object-Centric Representation Learning from Unlabeled Videos

Ruohan Gao, Dinesh Jayaraman, Kristen Grauman

cs.CV2023

The ObjectFolder Benchmark: Multisensory Learning with Neural and Real Objects

Ruohan Gao, Yiming Dou, Hao Li +5

cs.CV2025

Learning to Highlight Audio by Watching Movies

Chao Huang, Ruohan Gao, J. M. F. Tsang +5

cs.CV2023

An Extensible Multimodal Multi-task Object Dataset with Materials

Trevor Standley, Ruohan Gao, Dawn Chen +2

cs.RO2023

Differentiable Physics Simulation of Dynamics-Augmented Neural Objects

Simon Le Cleac'h, Hong-Xing Yu, Michelle Guo +5

cs.AI2026

Do Audio-Visual Large Language Models Really See and Hear?

Ramaneswaran Selvakumar, Kaousheik Jayakumar, S Sakshi +3

cs.CV2019

2.5D Visual Sound

Ruohan Gao, Kristen Grauman

physics.geo-ph2020

JULOC: A Local 3-D Refined Crust Model for the Geoneutrino Measurement at JUNO

Ruohan Gao, Zhiwei Li, Ran Han +7

cs.CV2025

ControlTac: Force- and Position-Controlled Tactile Data Augmentation with a Single Reference Image

Dongyu Luo, Kelin Yu, Amir-Hossein Shahidzadeh +3

cs.CV2020

VisualEchoes: Spatial Image Representation Learning through Echolocation

Ruohan Gao, Changan Chen, Ziad Al-Halah +2

cs.SD2024

SoundCam: A Dataset for Finding Humans Using Room Acoustics

Mason Wang, Samuel Clarke, Jui-Hsien Wang +2

cs.CV2023

Learning Object-Centric Neural Scattering Functions for Free-Viewpoint Relighting and Scene Composition

Hong-Xing Yu, Michelle Guo, Alireza Fathi +5

cs.RO2023

Sonicverse: A Multisensory Simulation Platform for Embodied Household Agents that See and Hear

Ruohan Gao, Hao Li, Gokul Dharan +6

cs.CV2018

Im2Flow: Motion Hallucination from Static Images for Action Recognition

Ruohan Gao, Bo Xiong, Kristen Grauman

eess.AS2025

Scene-wide Acoustic Parameter Estimation

Ricardo Falcon-Perez, Ruohan Gao, Gregor Mueckl +2

hep-ex2025

First measurement of reactor neutrino oscillations at JUNO

Angel Abusleme, Thomas Adam, Kai Adamowicz +1131

eess.AS2025

Aurelia: Test-time Reasoning Distillation in Audio-Visual LLMs

Sanjoy Chowdhury, Hanan Gani, Nishit Anand +5

cs.RO2023

NOIR: Neural Signal Operated Intelligent Robots for Everyday Activities

Ruohan Zhang, Sharon Lee, Minjune Hwang +11

cs.CV2022

Visual Acoustic Matching

Changan Chen, Ruohan Gao, Paul Calamia +1

cs.RO2021

ObjectFolder: A Dataset of Objects with Implicit Visual, Auditory, and Tactile Representations

Ruohan Gao, Yen-Yu Chang, Shivani Mall +2

cs.CV2021

VisualVoice: Audio-Visual Speech Separation with Cross-Modal Consistency

Ruohan Gao, Kristen Grauman

cs.RO2025

GenFlowRL: Shaping Rewards with Generative Object-Centric Flow in Visual Reinforcement Learning

Kelin Yu, Sheng Zhang, Harshit Soora +4

cs.CV2025

EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception

Sanjoy Chowdhury, Subrata Biswas, Sayan Nag +7

cs.SD2024

Hearing Anything Anywhere

Mason Wang, Ryosuke Sawata, Samuel Clarke +3

cs.DC2025

FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees

Gabriele Oliaro, Xupeng Miao, Xinhao Cheng +9

cs.RO2026

OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies

Kelin Yu, Haode Zhang, Harish Ravichandar +2

cs.SD2024

DiffSound: Differentiable Modal Sound Rendering and Inverse Rendering for Diverse Inference Tasks

Xutong Jin, Chenxi Xu, Ruohan Gao +3

cs.CV2018

ShapeCodes: Self-Supervised Feature Learning by Lifting Views to Viewgrids

Dinesh Jayaraman, Ruohan Gao, Kristen Grauman

cs.CV2019

Co-Separating Sounds of Visual Objects

Ruohan Gao, Kristen Grauman

cs.CV2021

Learning to Set Waypoints for Audio-Visual Navigation

Changan Chen, Sagnik Majumder, Ziad Al-Halah +3

hep-ex2025

Prospects for geoneutrino detection with JUNO

Thomas Adam, Shakeel Ahmad, Rizwan Ahmed +625

cs.CV2024

Spherical World-Locking for Audio-Visual Localization in Egocentric Videos

Heeseung Yun, Ruohan Gao, Ishwarya Ananthabhotla +6

eess.AS2025

Towards Perception-Informed Latent HRTF Representations

You Zhang, Andrew Francl, Ruohan Gao +3

cs.RO2026

HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos

Zhi Wang, Botao He, Kelin Yu +4

physics.geo-ph2024

Expected geoneutrino signal at JUNO using local integrated 3-D refined crustal model

Ran Han, ZhiWei Li, Ruohan Gao +11

cs.CV2024

Meerkat: Audio-Visual Large Language Model for Grounding in Space and Time

Sanjoy Chowdhury, Sayan Nag, Subhrajyoti Dasgupta +4

cs.CV2025

Differentiable Room Acoustic Rendering with Multi-View Vision Priors

Derong Jin, Ruohan Gao

cs.CV2024

The Audio-Visual Conversational Graph: From an Egocentric-Exocentric Perspective

Wenqi Jia, Miao Liu, Hao Jiang +4

cs.RO2022

See, Hear, and Feel: Smart Sensory Fusion for Robotic Manipulation

Hao Li, Yizhi Zhang, Junzhe Zhu +7

cs.CV2018

Learning to Separate Object Sounds by Watching Unlabeled Video

Ruohan Gao, Rogerio Feris, Kristen Grauman

cs.CV2025

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs

Sanjoy Chowdhury, Sayan Nag, Subhrajyoti Dasgupta +4

cs.SD2023

RealImpact: A Dataset of Impact Sound Fields for Real Objects

Samuel Clarke, Ruohan Gao, Mason Wang +5