papers

Publications (27)

cs.CV2024

Correspondence-Free SE(3) Point Cloud Registration in RKHS via Unsupervised Equivariant Learning

Ray Zhang, Zheming Zhou, Min Sun +5

This paper introduces a robust unsupervised SE(3) point cloud registration method that operates without requiring point correspondences. The method frames point clouds as functions…

eess.SY2026

Dynamic Association of Semantics and Parameter Estimates by Filtering

Marcus Greiff, Ray Zhang, Thomas Lew +1

We propose a probabilistic semantic filtering framework in which parameters of a dynamical system are inferred and associated with a closed set of semantic classes in a map. We ext…

cs.CV2026

Generalized-CVO: Fast and Correspondence-Free Local Point Cloud Registration with Second Order Riemannian Optimization

Ray Zhang, Marcus Greiff, Thomas Lew +1

We propose a fast and correspondence-free local point cloud registration method that leverages geometric surface structure and reproducing kernel Hilbert space (RKHS) embeddings. T…

cs.RO2024

RKHS-BA: A Robust Correspondence-Free Multi-View Registration Framework with Semantic Point Clouds

Ray Zhang, Jingwei Song, Xiang Gao +5

This work reports a novel multi-frame Bundle Adjustment (BA) framework called RKHS-BA. It uses continuous landmark representations that encode RGB-D/LiDAR and semantic observations…

cs.CV2024

Any2Point: Empowering Any-modality Large Models for Efficient 3D Understanding

Yiwen Tang, Ray Zhang, Jiaming Liu +8

Large foundation models have recently emerged as a prominent focus of interest, attaining superior performance in widespread scenarios. Due to the scarcity of 3D data, many efforts…

cs.CV2023

FM-OV3D: Foundation Model-based Cross-modal Knowledge Blending for Open-Vocabulary 3D Detection

Dongmei Zhang, Chang Li, Ray Zhang +4

The superior performances of pre-trained foundation models in various visual tasks underscore their potential to enhance the 2D models' open-vocabulary ability. Existing methods ex…

cs.CV2025

SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems

Ziyu Guo, Ray Zhang, Hao Chen +4

The rapid advancement of Large Multi-modal Models (LMMs) has enabled their application in scientific problem-solving, yet their fine-grained capabilities remain under-explored. In…

cs.SE2025

What Types of Code Review Comments Do Developers Most Frequently Resolve?

Saul Goldman, Hong Yi Lin, Jirat Pasuksmit +11

Large language model (LLM)-powered code review automation tools have been introduced to generate code review comments. However, not all generated comments will drive code changes.…

cs.RO2024

SLAM assisted 3D tracking system for laparoscopic surgery

Jingwei Song, Ray Zhang, Wenwei Zhang +2

A major limitation of minimally invasive surgery is the difficulty in accurately locating the internal anatomical structures of the target organ due to the lack of tactile feedback…

cs.CV2025

CrossLMM: Decoupling Long Video Sequences from LMMs via Dual Cross-Attention Mechanisms

Shilin Yan, Jiaming Han, Joey Tsai +5

The advent of Large Multimodal Models (LMMs) has significantly enhanced Large Language Models (LLMs) to process and interpret diverse data modalities (e.g., image and video). Howev…

cs.CV2025

Exploring the Potential of Encoder-free Architectures in 3D LMMs

Yiwen Tang, Zoey Guo, Zhuhao Wang +8

Encoder-free architectures have been preliminarily explored in the 2D Large Multimodal Models (LMMs), yet it remains an open question whether they can be effectively applied to 3D…

cs.CV2024

MR-MLLM: Mutual Reinforcement of Multimodal Comprehension and Vision Perception

Guanqun Wang, Xinyu Wei, Jiaming Liu +5

In recent years, multimodal large language models (MLLMs) have shown remarkable capabilities in tasks like visual question answering and common sense reasoning, while visual percep…

q-bio.QM2021

Stain-free Detection of Embryo Polarization using Deep Learning

Cheng Shen, Adiyant Lamba, Meng Zhu +3

Polarization of the mammalian embryo at the right developmental time is critical for its development to term and would be valuable in assessing the potential of human embryos. Howe…

cs.CV2025

Are We Ready for RL in Text-to-3D Generation? A Progressive Investigation

Yiwen Tang, Zoey Guo, Kaixin Zhu +11

Reinforcement learning (RL), earlier proven to be effective in large language and multi-modal models, has been successfully extended to enhance 2D image generation recently. Howeve…

cs.CV2026

InterleaveThinker: Reinforcing Agentic Interleaved Generation

Dian Zheng, Harry Lee, Manyuan Zhang +4

Recent image generators have demonstrated impressive photorealism and instruction-following capabilities in single-image generation and editing. However, constrained by their archi…

cs.RO2021

Legged Robot State Estimation using Invariant Kalman Filtering and Learned Contact Events

Tzu-Yuan Lin, Ray Zhang, Justin Yu +1

This work develops a learning-based contact estimator for legged robots that bypasses the need for physical sensors and takes multi-modal proprioceptive sensory data as input. Unli…

astro-ph.CO2009

Disks in the sky: A reassessment of the WMAP "cold spot"

Ray Zhang, Dragan Huterer

We reassess the evidence that WMAP temperature maps contain a statistically significant "cold spot" by repeating the analysis using simple circular top-hat (disk) weights, as well…

eess.SY2025

Semantic Property Maps for Driving Applications

Marcus Greiff, Ray Zhang, Takeru Shirasawa +1

We consider the problem of estimating the parameters of a vehicle dynamics model for predictive control in driving applications. Instead of solely using the instantaneous parameter…

cs.CV2023

Cloud-Device Collaborative Learning for Multimodal Large Language Models

Guanqun Wang, Jiaming Liu, Chenxuan Li +8

The burgeoning field of Multimodal Large Language Models (MLLMs) has exhibited remarkable performance in diverse tasks such as captioning, commonsense reasoning, and visual scene u…

cs.RO2020

Bayesian Spatial Kernel Smoothing for Scalable Dense Semantic Mapping

Lu Gan, Ray Zhang, Jessy W. Grizzle +2

This paper develops a Bayesian continuous 3D semantic occupancy map from noisy point clouds by generalizing the Bayesian kernel inference model for building occupancy maps, a binar…

cs.CV2023

LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR Understanding

Senqiao Yang, Jiaming Liu, Ray Zhang +7

Recently, Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) have shown promise in instruction following and 2D image understanding. While these models are p…

cs.CV2026

Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning

Hohin Kwan, Hongyu Li, Ray Zhang +5

Recent interest in multimodal large language models (MLLMs) raises a central question: can they reason over dynamic visual evidence rather than merely recognize objects or events i…

cs.RO2026

Vision-Conditioned Variational Bayesian Last Layer Dynamics Models

Paul Brunzema, Thomas Lew, Ray Zhang +3

Agile control of robotic systems often requires anticipating how the environment affects system behavior. For example, a driver must perceive the road ahead to anticipate available…

cs.CV2024

Point-PEFT: Parameter-Efficient Fine-Tuning for 3D Pre-trained Models

Yiwen Tang, Ray Zhang, Zoey Guo +4

The popularity of pre-trained large models has revolutionized downstream tasks across diverse fields, such as language, vision, and multi-modality. To minimize the adaption cost fo…

cs.RO2023

BDIS-SLAM: A lightweight CPU-based dense stereo SLAM for surgery

Jingwei Song, Ray Zhang, Qiuchen Zhu +2

Purpose: Common dense stereo Simultaneous Localization and Mapping (SLAM) approaches in Minimally Invasive Surgery (MIS) require high-end parallel computational resources for real-…

cs.CV2020

A New Framework for Registration of Semantic Point Clouds from Stereo and RGB-D Cameras

Ray Zhang, Tzu-Yuan Lin, Chien Erh Lin +5

This paper reports on a novel nonparametric rigid point cloud registration framework that jointly integrates geometric and semantic measurements such as color or semantic labels in…

cs.CV2023

ViewRefer: Grasp the Multi-view Knowledge for 3D Visual Grounding with GPT and Prototype Guidance

Zoey Guo, Yiwen Tang, Ray Zhang +4

Understanding 3D scenes from multi-view inputs has been proven to alleviate the view discrepancy issue in 3D visual grounding. However, existing methods normally neglect the view c…