2 citations · 4 across the 3 of their papers we have counts for
4 papers · 1 filter
SSR: Pushing the Limit of Spatial Intelligence with Structured Scene Reasoning
Yi Zhang, Youya Xia, Yong Wang +7
While Multimodal Large Language Models (MLLMs) excel in semantic tasks, they frequently lack the "spatial sense" essential for sophisticated geometric reasoning. Current models typ…
Image-to-Image Translation for Autonomous Driving from Coarsely-Aligned Image Pairs
Youya Xia, Josephine Monica, Wei-Lun Chao +3
A self-driving car must be able to reliably handle adverse weather conditions (e.g., snowy) to operate safely. In this paper, we investigate the idea of turning sensor inputs (i.e.…
Fast Underwater Image Enhancement for Improved Visual Perception
Md Jahidul Islam, Youya Xia, Junaed Sattar
In this paper, we present a conditional generative adversarial network-based model for real-time underwater image enhancement. To supervise the adversarial training, we formulate a…
Visual Diver Recognition for Underwater Human-Robot Collaboration
Youya Xia, Junaed Sattar
This paper presents an approach for autonomous underwater robots to visually detect and identify divers. The proposed approach enables an autonomous underwater robot to detect mult…