Publications (36)
Matryoshka Concept Bottleneck Models
Ziye Chen, Hongbin Lin, Jie Li +1
Concept Bottleneck Models (CBMs) have emerged as a prominent paradigm for interpretable deep learning, learning by grounding predictions in human-understandable concepts. However,…
Source-free Domain Adaptation via Avatar Prototype Generation and Adaptation
Zhen Qiu, Yifan Zhang, Hongbin Lin +4
We study a practical domain adaptation task, called source-free unsupervised domain adaptation (UDA) problem, in which we cannot access source domain data due to data privacy issue…
Open-source High-precision Autonomous Suturing Framework With Visual Guidance
Hongbin Lin, Bin Li, Yunhui Liu +1
Autonomous surgery has attracted increasing attention for revolutionizing robotic patient care, yet remains a distant and challenging goal. In this paper, we propose an image-based…
Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent
Yao Shu, Chenxing Wei, Hongbin Lin +2
Online reinforcement learning with verifiable rewards (RLVR) turns checkable outcomes into a scalable training signal, but it keeps rollout generation, verifier scoring, and refere…
AR1-ZO: Topology-Aware Rank-1 Zeroth-Order Queries for High-Rank LoRA Fine-Tuning
Ziye Chen, Hongbin Lin, Chenyu Zhang +3
Zeroth-order (ZO) optimization enables large-language-model fine-tuning without storing backpropagation activations, while LoRA supplies compact trainable adapters. Combining them…
DriveFlow: Rectified Flow Adaptation for Robust 3D Object Detection in Autonomous Driving
Hongbin Lin, Yiming Yang, Chaoda Zheng +7
In autonomous driving, vision-centric 3D object detection recognizes and localizes 3D objects from RGB images. However, due to high annotation costs and diverse outdoor scenes, tra…
Visuomotor Grasping with World Models for Surgical Robots
Hongbin Lin, Bin Li, Kwok Wai Samuel Au
Grasping is a fundamental task in robot-assisted surgery (RAS), and automating it can reduce surgeon workload while enhancing efficiency, safety, and consistency beyond teleoperate…
Towards Multi-dimensional Explanation Alignment for Medical Classification
Lijie Hu, Songning Lai, Wenshuo Chen +5
The lack of interpretability in the field of medical image analysis has significant ethical and legal implications. Existing interpretable methods in this domain encounter several…
UNeR3D: Versatile and Scalable 3D RGB Point Cloud Generation from 2D Images in Unsupervised Reconstruction
Hongbin Lin, Juangui Xu, Qingfeng Xu +5
In the realm of 3D reconstruction from 2D images, a persisting challenge is to achieve high-precision reconstructions devoid of 3D Ground Truth data reliance. We present UNeR3D, a…
End-to-End Learning of Deep Visuomotor Policy for Needle Picking
Hongbin Lin, Bin Li, Xiangyu Chu +3
Needle picking is a challenging manipulation task in robot-assisted surgery due to the characteristics of small slender shapes of needles, needles' variations in shapes and sizes,…
Fast, Robust, and Versatile Event Detection through HMM Belief State Gradient Measures
Shuangqi Luo, Hongmin Wu, Hongbin Lin +3
Event detection is a critical feature in data-driven systems as it assists with the identification of nominal and anomalous behavior. Event detection is increasingly relevant in ro…
Robot Introspection with Bayesian Nonparametric Vector Autoregressive Hidden Markov Models
Hongmin Wu, Hongbin Lin, Yisheng Guan +2
Robot introspection, as opposed to anomaly detection typical in process monitoring, helps a robot understand what it is doing at all times. A robot should be able to identify its a…
DriveGEN: Generalized and Robust 3D Detection in Driving via Controllable Text-to-Image Diffusion Generation
Hongbin Lin, Zilu Guo, Yifan Zhang +5
In autonomous driving, vision-centric 3D detection aims to identify 3D objects from images. However, high data collection costs and diverse real-world scenarios limit the scale of…
Benchmarking and Mitigating Sycophancy in Medical Vision Language Models
Juangui Xu, Zikun Guo, Jingwei Lv +5
Visual language models (VLMs) have the potential to transform medical workflows. However, the deployment is limited by sycophancy. Despite this serious threat to patient safety, a…
Retrieval-Infused Reasoning Sandbox: A Benchmark for Decoupling Retrieval and Reasoning Capabilities
Shuangshuang Ying, Zheyu Wang, Yunjian Peng +16
Despite strong performance on existing benchmarks, it remains unclear whether large language models can reason over genuinely novel scientific information. Most evaluations score e…
Algorithmic Recourse of In-Context Learning for Tabular Data
Wenshuo Dong, Jiaming Zhang, Shaopeng Fu +3
The paper introduces a theoretical and practical framework for providing algorithmic recourse on tabular data using in-context learning with large language models, proposing a zero…
EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments
Deyao Zhu, Xin Zhou, Shengling Qin +44
Pretraining scaling laws reveal that model capability improves predictably with data and compute. But learning from real world environments after deployment remains far less unders…
A Reliable Gravity Compensation Control Strategy for dVRK Robotic Arms With Nonlinear Disturbance Forces
Hongbin Lin, C. W. Vincent Hui, Yan Wang +3
External disturbance forces caused by nonlinear springy electrical cables in the Master Tool Manipulator (MTM) of the da Vinci Research Kit (dVRK) limits the usage of the existing…
Multi-Group Equivariant Augmentation for Reinforcement Learning in Robot Manipulation
Hongbin Lin, Juan Rojas, Kwok Wai Samuel Au
Sampling efficiency is critical for deploying visuomotor learning in real-world robotic manipulation. While task symmetry has emerged as a promising inductive bias to improve effic…
Learning Deep Nets for Gravitational Dynamics with Unknown Disturbance through Physical Knowledge Distillation: Initial Feasibility Study
Hongbin Lin, Qian Gao, Xiangyu Chu +4
Learning high-performance deep neural networks for dynamic modeling of high Degree-Of-Freedom (DOF) robots remains challenging due to the sampling complexity. Typical unknown syste…
TopoStreamer: Temporal Lane Segment Topology Reasoning in Autonomous Driving
Yiming Yang, Yueru Luo, Bingkun He +8
Lane segment topology reasoning constructs a comprehensive road network by capturing the topological relationships between lane segments and their semantic types. This enables end-…
X-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End Driving
Chaoda Zheng, Sean Li, Jinhao Deng +9
Scalable and reliable evaluation is increasingly critical in the end-to-end era of autonomous driving, where vision--language--action (VLA) policies directly map raw sensor streams…
Graph-Guided Dual-Level Augmentation for 3D Scene Segmentation
Hongbin Lin, Yifan Jiang, Juangui Xu +5
3D point cloud segmentation aims to assign semantic labels to individual points in a scene for fine-grained spatial understanding. Existing methods typically adopt data augmentatio…
Fully Test-Time Adaptation for Monocular 3D Object Detection
Hongbin Lin, Yifan Zhang, Shuaicheng Niu +2
Monocular 3D object detection (Mono 3Det) aims to identify 3D objects from a single RGB image. However, existing methods often assume training and test data follow the same distrib…
Controllable Concept Bottleneck Models
Hongbin Lin, Chenyang Ren, Juangui Xu +7
Concept Bottleneck Models (CBMs) have garnered much attention for their ability to elucidate the prediction process through a human-understandable concept layer. However, most prev…
Online Robot Introspection via Wrench-based Action Grammars
Juan Rojas, Shuangqi Luo, Dingqiao Zhu +5
Robotic failure is all too common in unstructured robot tasks. Despite well-designed controllers, robots often fail due to unexpected events. How do robots measure unexpected event…
Imbalance-Agnostic Source-Free Domain Adaptation via Avatar Prototype Alignment
Hongbin Lin, Mingkui Tan, Yifan Zhang +5
Source-free Unsupervised Domain Adaptation (SF-UDA) aims to adapt a well-trained source model to an unlabeled target domain without access to the source data. One key challenge is…
FASTopoWM: Fast-Slow Lane Segment Topology Reasoning with Latent World Models
Yiming Yang, Hongbin Lin, Yueru Luo +7
Lane segment topology reasoning provides comprehensive bird's-eye view (BEV) road scene understanding, which can serve as a key perception module in planning-oriented end-to-end au…
FutureX: Enhance End-to-End Autonomous Driving via Latent Chain-of-Thought World Model
Hongbin Lin, Yiming Yang, Yifan Zhang +10
In autonomous driving, end-to-end planners learn scene representations from raw sensor data and utilize them to generate a motion plan or control actions. However, exclusive relian…
Prototype-Guided Continual Adaptation for Class-Incremental Unsupervised Domain Adaptation
Hongbin Lin, Yifan Zhang, Zhen Qiu +4
This paper studies a new, practical but challenging problem, called Class-Incremental Unsupervised Domain Adaptation (CI-UDA), where the labeled source domain contains all classes,…
Recovering from External Disturbances in Online Manipulation through State-Dependent Revertive Recovery Policies
Hongmin Wu, Hongbin Lin, Shuangqi Luo +3
Robots are increasingly entering uncertain and unstructured environments. Within these, robots are bound to face unexpected external disturbances like accidental human or tool coll…
NoiseRater: Meta-Learned Noise Valuation for Diffusion Model Training
Fang Wu, Haokai Zhao, Da Xing +17
Diffusion models have achieved remarkable success across a wide range of generative tasks, yet their training paradigm largely treats injected noise as uniformly informative. In th…
Editable Concept Bottleneck Models
Lijie Hu, Chenyang Ren, Zhengyu Hu +5
Concept Bottleneck Models (CBMs) have garnered much attention for their ability to elucidate the prediction process through a humanunderstandable concept layer. However, most previ…
World Models for General Surgical Grasping
Hongbin Lin, Bin Li, Chun Wai Wong +3
Intelligent vision control systems for surgical robots should adapt to unknown and diverse objects while being robust to system disturbances. Previous methods did not meet these re…
SSIM-Variation-Based Complexity Optimization for Versatile Video Coding
Jielian Lin, Hongbin Lin, Zhichen Zhang +2
To date, Versatile Video Coding (VVC) has a more magnificent overall performance than High Efficiency Video Coding (HEVC). The Quadtree with Nested Multi-Type Tree (QTMT) coding bl…
PiSA: A Self-Augmented Data Engine and Training Strategy for 3D Understanding with Large Models
Zilu Guo, Hongbin Lin, Zhihao Yuan +6
3D Multimodal Large Language Models (MLLMs) have recently made substantial advancements. However, their potential remains untapped, primarily due to the limited quantity and subopt…