activity
20222024
most citedControllable Video Generation by Learning the Underlying Dynamical System with Neural ODE

1 citations · 2 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CL2024

Audio Dialogues: Dialogues dataset for audio and music understanding

Arushi Goel, Zhifeng Kong, Rafael Valle +1

Existing datasets for audio understanding primarily focus on single-turn interactions (i.e. audio captioning, audio question answering) for describing audio in natural language, th…

cs.RO20231 cited

Language-guided Robot Grasping: CLIP-based Referring Grasp Synthesis in Clutter

Georgios Tziafas, Yucheng Xu, Arushi Goel +3

Robots operating in human-centric environments require the integration of visual grounding and grasping capabilities to effectively manipulate objects based on user instructions. T…

cs.CL2023

Semi-supervised multimodal coreference resolution in image narrations

Arushi Goel, Basura Fernando, Frank Keller +1

In this paper, we study multimodal coreference resolution, specifically where a longer descriptive text, i.e., a narration is paired with an image. This poses significant challenge…

cs.CV2023

Encyclopedic VQA: Visual questions about detailed properties of fine-grained categories

Thomas Mensink, Jasper Uijlings, Lluis Castrejon +6

We propose Encyclopedic-VQA, a large scale visual question answering (VQA) dataset featuring visual questions about detailed properties of fine-grained categories and instances. It…

cs.CV20231 cited

Controllable Video Generation by Learning the Underlying Dynamical System with Neural ODE

Yucheng Xu, Li Nanbo, Arushi Goel +5

Videos depict the change of complex dynamical systems over time in the form of discrete image sequences. Generating controllable videos by learning the dynamical system is an impor…

cs.CV2022

WiCV 2022: The Tenth Women In Computer Vision Workshop

Doris Antensteiner, Silvia Bucci, Arushi Goel +6

In this paper, we present the details of Women in Computer Vision Workshop - WiCV 2022, organized alongside the hybrid CVPR 2022 in New Orleans, Louisiana. It provides a voice to a…