papers

Publications (31)

cs.CV2026

Vision Also You Need: Navigating Out-of-Distribution Detection with Multimodal Large Language Model

Haoran Xu, Yanlin Liu, Zizhao Tong +8

Out-of-Distribution (OOD) detection is a critical task that has garnered significant attention. The emergence of CLIP has spurred extensive research into zero-shot OOD detection, o…

cs.CV2024

Separate and Conquer: Decoupling Co-occurrence via Decomposition and Representation for Weakly Supervised Semantic Segmentation

Zhiwei Yang, Kexue Fu, Minghong Duan +3

Weakly supervised semantic segmentation (WSSS) with image-level labels aims to achieve segmentation tasks without dense annotations. However, attributed to the frequent coupling of…

cs.CV2022

DSPoint: Dual-scale Point Cloud Recognition with High-frequency Fusion

Renrui Zhang, Ziyao Zeng, Ziyu Guo +3

Point cloud processing is a challenging task due to its sparsity and irregularity. Prior works introduce delicate designs on either local feature aggregator or global geometric arc…

cs.CV2022

Robust Point Cloud Registration Framework Based on Deep Graph Matching(TPAMI Version)

Kexue Fu, Jiazheng Luo, Xiaoyuan Luo +3

3D point cloud registration is a fundamental problem in computer vision and robotics. Recently, learning-based point cloud registration methods have made great progress. However, t…

cs.CV2026

Region Matters: Efficient and Reliable Region-Aware Visual Place Recognition

Shunpeng Chen, Yukun Song, Changwei Wang +6

Visual Place Recognition (VPR) determines a query image's geographic location by matching it against geotagged databases. However, existing methods struggle with perceptual aliasin…

cs.CV2025

CurriFlow: Curriculum-Guided Depth Fusion with Optical Flow-Based Temporal Alignment for 3D Semantic Scene Completion

Jinzhou Lin, Jie Zhou, Wenhao Xu +7

Semantic Scene Completion (SSC) aims to infer complete 3D geometry and semantics from monocular images, serving as a crucial capability for camera-based perception in autonomous dr…

cs.CV2022

Distillation with Contrast is All You Need for Self-Supervised Point Cloud Representation Learning

Kexue Fu, Peng Gao, Renrui Zhang +3

In this paper, we propose a simple and general framework for self-supervised point cloud representation learning. Human beings understand the 3D world by extracting two levels of i…

cs.AI2025

Large Language Model-Based Intelligent Antenna Design System

Tao Wu, Kexue Fu, Qiang Hua +2

Antenna simulation typically involves modeling and optimization, which are time-consuming and labor-intensive, slowing down antenna analysis and design. This paper presents a proto…

cs.HC2026

Spatial Balancing: Designing an LLM-Powered Spatial Externalization Interface for Iterative Science Communication Writing

Kexue Fu, Jiaye Leng, Yawen Zhang +7

Science communication revision requires writers to dynamically balance scientific exposition and narrative engagement - a process where writers often struggle with competing direct…

cs.CV2025

Focus on Local: Finding Reliable Discriminative Regions for Visual Place Recognition

Changwei Wang, Shunpeng Chen, Yukun Song +11

Visual Place Recognition (VPR) is aimed at predicting the location of a query image by referencing a database of geotagged images. For VPR task, often fewer discriminative local re…

cs.HC2025

RetroChat: Designing for the Preservation of Past Digital Experiences

Suifang Zhou, Kexue Fu, Huanmin Yi +1

Rapid changes in social networks have transformed the way people express themselves, turning past neologisms, values, and mindsets embedded in these expressions into online heritag…

cs.CV2025

Exploring CLIP's Dense Knowledge for Weakly Supervised Semantic Segmentation

Zhiwei Yang, Yucong Meng, Kexue Fu +3

Weakly Supervised Semantic Segmentation (WSSS) with image-level labels aims to achieve pixel-level predictions using Class Activation Maps (CAMs). Recently, Contrastive Language-Im…

cs.CV2025

MoRe: Class Patch Attention Needs Regularization for Weakly Supervised Semantic Segmentation

Zhiwei Yang, Yucong Meng, Kexue Fu +2

Weakly Supervised Semantic Segmentation (WSSS) with image-level labels typically uses Class Activation Maps (CAM) to achieve dense predictions. Recently, Vision Transformer (ViT) h…

cs.CV2022

Boosting Point-BERT by Multi-choice Tokens

Kexue Fu, Mingzhi Yuan, Manning Wang

Masked language modeling (MLM) has become one of the most successful self-supervised pre-training task. Inspired by its success, Point-BERT, as a pioneer work in point cloud, propo…

cs.CV2024

The Rise of AI Language Pathologists: Exploring Two-level Prompt Learning for Few-shot Weakly-supervised Whole Slide Image Classification

Linhao Qu, Xiaoyuan Luo, Kexue Fu +2

This paper introduces the novel concept of few-shot weakly supervised learning for pathology Whole Slide Image (WSI) classification, denoted as FSWC. A solution is proposed based o…

cs.CV2026

DiCLIP: Diffusion Model Enhances CLIP's Dense Knowledge for Weakly Supervised Semantic Segmentation

Zhiwei Yang, Pengfei Song, Yucong Meng +3

Weakly Supervised Semantic Segmentation (WSSS) with image-level labels typically leverages Class Activation Maps (CAMs) to achieve pixel-level predictions. Recently, Contrastive La…

cs.HC2026

Vistoria: A Multimodal System to Support Fictional Story Writing through Instrumental Text-Image Co-Editing

Kexue Fu, Jingfei Huang, Long Ling +4

Humans think visually-we remember in images, dream in pictures, and use visual metaphors to communicate. Yet, most creative writing tools remain text-centric, limiting how authors…

cs.CV2026

Adaptive Visual Autoregressive Acceleration via Dual-Linkage Entropy Analysis

Yu Zhang, Jingyi Liu, Feng Liu +5

Visual AutoRegressive modeling (VAR) suffers from substantial computational cost due to the massive token count involved. Failing to account for the continuous evolution of modelin…

cs.HC2025

RevTogether: Supporting Science Story Revision with Multiple AI Agents

Yu Zhang, Kexue Fu, Zhicong Lu

As a popular form of science communication, science stories attract readers because they combine engaging narratives with comprehensible scientific knowledge. However, crafting suc…

cs.CV2024

FAST: A Dual-tier Few-Shot Learning Paradigm for Whole Slide Image Classification

Kexue Fu, Xiaoyuan Luo, Linhao Qu +5

The expensive fine-grained annotation and data scarcity have become the primary obstacles for the widespread adoption of deep learning-based Whole Slide Images (WSI) classification…

cs.CV2024

Tackling Ambiguity from Perspective of Uncertainty Inference and Affinity Diversification for Weakly Supervised Semantic Segmentation

Zhiwei Yang, Yucong Meng, Kexue Fu +2

Weakly supervised semantic segmentation (WSSS) with image-level labels intends to achieve dense tasks without laborious annotations. However, due to the ambiguous contexts and fuzz…

cs.RO2025

A Constructed Response: Designing and Choreographing Robot Arm Movements in Collaborative Dance Improvisation

Xiaoyu Chang, Fan Zhang, Kexue Fu +3

Dancers often prototype movements themselves or with each other during improvisation and choreography. How are these interactions altered when physically manipulable technologies a…

cs.HC2025

From Temporal to Spatial: Designing Spatialized Interactions with Segmented-audios in Immersive Environments for Active Engagement with Performing Arts Intangible Cultural Heritage

Yuqi Wang, Sirui Wang, Shiman Zhang +3

Performance artforms like Peking opera face transmission challenges due to the extensive passive listening required to understand their nuance. To create engaging forms of experien…

cs.HC2025

"Becoming My Own Audience": How Dancers React to Avatars Unlike Themselves in Motion Capture-Supported Live Improvisational Performance

Fan Zhang, Molin Li, Xiaoyu Chang +3

The use of motion capture in live dance performances has created an emerging discipline enabling dancers to play different avatars on the digital stage. Unlike classical workflows,…

cs.CV2021

Robust Point Cloud Registration Framework Based on Deep Graph Matching

Kexue Fu, Shaolei Liu, Xiaoyuan Luo +1

3D point cloud registration is a fundamental problem in computer vision and robotics. There has been extensive research in this area, but existing methods meet great challenges in…

cs.CV2023

PointMBF: A Multi-scale Bidirectional Fusion Network for Unsupervised RGB-D Point Cloud Registration

Mingzhi Yuan, Kexue Fu, Zhihao Li +2

Point cloud registration is a task to estimate the rigid transformation between two unaligned scans, which plays an important role in many computer vision applications. Previous le…

cs.RO2023

Boosting 3D Point Cloud Registration by Transferring Multi-modality Knowledge

Mingzhi Yuan, Xiaoshui Huang, Kexue Fu +2

The recent multi-modality models have achieved great performance in many vision tasks because the extracted features contain the multi-modality knowledge. However, most of the curr…

cs.CV2022

POS-BERT: Point Cloud One-Stage BERT Pre-Training

Kexue Fu, Peng Gao, ShaoLei Liu +3

Recently, the pre-training paradigm combining Transformer and masked language modeling has achieved tremendous success in NLP, images, and point clouds, such as BERT. However, dire…

cs.HC2026

The Configuration of Space: Probing the Way Social Interaction and Perception are Affected by Task-Specific Spatial Representations in Online Video Communication

Yihuan Chen, Kexue Fu, Qianyi Chen +2

Humans live and act in 3D space, but often work and communicate on 2D surfaces. The prevalence of online communication on 2D screens raises the issue of whether human spatial confi…

cs.CV2021

Multi-View Partial (MVP) Point Cloud Challenge 2021 on Completion and Registration: Methods and Results

Liang Pan, Tong Wu, Zhongang Cai +26

As real-scanned point clouds are mostly partial due to occlusions and viewpoints, reconstructing complete 3D shapes based on incomplete observations becomes a fundamental problem f…

cs.CV2021

A Learnable Self-supervised Task for Unsupervised Domain Adaptation on Point Clouds

Xiaoyuan Luo, Shaolei Liu, Kexue Fu +2

Deep neural networks have achieved promising performance in supervised point cloud applications, but manual annotation is extremely expensive and time-consuming in supervised learn…