papers

Publications (16)

cs.CV2026

Geometry-Guided 3D Visual Token Pruning for Video-Language Models

Han Li, Zehao Huang, Jiahui Fu +2

Multimodal large language models have demonstrated remarkable capabilities in 2D vision, motivating their extension to 3D scene understanding. Recent studies represent 3D scenes as…

cs.CV2026

ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?

Han Li, Si Liu, Zehao Huang +6

Multimodal Large Language Models (MLLMs) have achieved remarkable success across diverse expert-level tasks, but they still struggle with fundamental abilities that humans naturall…

cs.RO2021

A Multi-Hypothesis Approach to Pose Ambiguity in Object-Based SLAM

Jiahui Fu, Qiangqiang Huang, Kevin Doherty +2

In object-based Simultaneous Localization and Mapping (SLAM), 6D object poses offer a compact representation of landmark geometry useful for downstream planning and manipulation ta…

cs.CV2024

Eliminating Cross-modal Conflicts in BEV Space for LiDAR-Camera 3D Object Detection

Jiahui Fu, Chen Gao, Zitian Wang +4

Recent 3D object detectors typically utilize multi-sensor data and unify multi-modal features in the shared bird's-eye view (BEV) representation space. However, our empirical findi…

cs.CV2024

V2X-PC: Vehicle-to-everything Collaborative Perception via Point Cluster

Si Liu, Zihan Ding, Jiahui Fu +4

The objective of the collaborative vehicle-to-everything perception task is to enhance the individual vehicle's perception capability through message communication among neighborin…

cs.RO2022

PlaneSDF-based Change Detection for Long-term Dense Mapping

Jiahui Fu, Chengyuan Lin, Yuichi Taguchi +4

The ability to process environment maps across multiple sessions is critical for robots operating over extended periods of time. Specifically, it is desirable for autonomous agents…

eess.AS2026

Deep Hierarchical Knowledge Loss for Fault Intensity Diagnosis

Yu Sha, Shuiping Gou, Bo Liu +8

Fault intensity diagnosis (FID) plays a pivotal role in intelligent manufacturing while neglecting dependencies among target classes hinders its practical deployment. This paper in…

cs.RO2025

NovaFlow: Zero-Shot Manipulation via Actionable Flow from Generated Videos

Hongyu Li, Lingfeng Sun, Yafei Hu +4

Enabling robots to execute novel manipulation tasks zero-shot is a central goal in robotics. Most existing methods assume in-distribution tasks or rely on fine-tuning with embodime…

cs.CV2022

3D-SPS: Single-Stage 3D Visual Grounding via Referred Point Progressive Selection

Junyu Luo, Jiahui Fu, Xianghao Kong +5

3D visual grounding aims to locate the referred target object in 3D point cloud scenes according to a free-form language description. Previous methods mostly follow a two-stage par…

cs.CV2023

Object as Query: Lifting any 2D Object Detector to 3D Detection

Zitian Wang, Zehao Huang, Jiahui Fu +2

3D object detection from multi-view images has drawn much attention over the past few years. Existing methods mainly establish 3D representations from multi-view images and adopt a…

cs.CV2026

Show, Don't Tell: Detecting Novel Objects by Watching Human Videos

James Akl, Jose Nicolas Avendano Arbelaez, James Barabas +16

How can a robot quickly identify and recognize new objects shown to it during a human demonstration? Existing closed-set object detectors frequently fail at this because the object…

cs.RO2026

NovaPlan: Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning

Jiahui Fu, Junyu Nan, Lingfeng Sun +5

Solving long-horizon tasks requires robots to integrate high-level semantic reasoning with low-level physical interaction. While vision-language models (VLMs) and video generation…

cs.CV2021

Improved Pillar with Fine-grained Feature for 3D Object Detection

Jiahui Fu, Guanghui Ren, Yunpeng Chen +1

3D object detection with LiDAR point clouds plays an important role in autonomous driving perception module that requires high speed, stability and accuracy. However, the existing…

cs.RO2023

NeuSE: Neural SE(3)-Equivariant Embedding for Consistent Spatial Understanding with Objects

Jiahui Fu, Yilun Du, Kurran Singh +2

We present NeuSE, a novel Neural SE(3)-Equivariant Embedding for objects, and illustrate how it supports object SLAM for consistent spatial understanding with long-term scene chang…

cs.CV2026

Generative Lane Topology Reasoning via Autoregressive Model with Geometry Prior

Jiahui Fu, Zehao Huang, Han Li +2

Lane topology reasoning aims to construct a lane graph from onboard sensor observations. Existing methods follow a detection and association paradigm that treats each lane instance…

cs.RO2022

Robust Change Detection Based on Neural Descriptor Fields

Jiahui Fu, Yilun Du, Kurran Singh +2

The ability to reason about changes in the environment is crucial for robots operating over extended periods of time. Agents are expected to capture changes during operation so tha…