papers

Publications (44)

cs.CV2021

A Compositional Feature Embedding and Similarity Metric for Ultra-Fine-Grained Visual Categorization

Yajie Sun, Miaohua Zhang, Xiaohan Yu +2

Fine-grained visual categorization (FGVC), which aims at classifying objects with small inter-class variances, has been significantly advanced in recent years. However, ultra-fine-…

cs.AI2026

MSearcher: Modular Multimodal Information Seeking Agency with Retrieval-Oriented Reasoning

Xiaohan Yu, Chao Feng, Lang Mei +1

Recent advances in DeepResearch-style agents have demonstrated strong capabilities in autonomous information acquisition and synthesize from real-world web environments. However, e…

cs.IR2025

CogPlanner: Unveiling the Potential of Agentic Multimodal Retrieval Augmented Generation with Planning

Xiaohan Yu, Zhihan Yang, Chong Chen

Multimodal Retrieval Augmented Generation (MRAG) systems have shown promise in enhancing the generation capabilities of multimodal large language models (MLLMs). However, existing…

cs.CV2026

Revisiting Photometric Ambiguity for Accurate Gaussian-Splatting Surface Reconstruction

Jiahe Li, Jiawei Zhang, Xiao Bai +4

Surface reconstruction with differentiable rendering has achieved impressive performance in recent years, yet the pervasive photometric ambiguities have strictly bottlenecked exist…

cs.CV2026

SCE-SLAM: Scale-Consistent Monocular SLAM via Scene Coordinate Embeddings

Yuchen Wu, Jiahe Li, Xiaohan Yu +3

Monocular visual SLAM enables 3D reconstruction from internet video and autonomous navigation on resource-constrained platforms, yet suffers from scale drift, i.e., the gradual div…

cs.IR2024

RA-Rec: An Efficient ID Representation Alignment Framework for LLM-based Recommendation

Xiaohan Yu, Li Zhang, Xin Zhao +2

Large language models (LLM) have recently emerged as a powerful tool for a variety of natural language processing tasks, bringing a new surge of combining LLM with recommendation s…

cs.RO2026

UGotMe: An Embodied System for Affective Human-Robot Interaction

Peizhen Li, Longbing Cao, Xiao-Ming Wu +2

Equipping humanoid robots with the capability to understand emotional states of human interactants and express emotions appropriately according to situations is essential for affec…

cs.CV2026

SpectralKAN: Weighted Activation Distribution Kolmogorov-Arnold Network for Hyperspectral Image Change Detection

Yanheng Wang, Xiaohan Yu, Yongsheng Gao +6

Kolmogorov-Arnold networks (KANs) represent data features by learning the activation functions and demonstrate superior accuracy with fewer parameters, FLOPs, GPU memory usage (Mem…

cs.CV2025

Visual and Semantic Prompt Collaboration for Generalized Zero-Shot Learning

Huajie Jiang, Zhengxian Li, Xiaohan Yu +4

Generalized zero-shot learning aims to recognize both seen and unseen classes with the help of semantic information that is shared among different classes. It inevitably requires c…

cs.CV2025

SparseSurf: Sparse-View 3D Gaussian Splatting for Surface Reconstruction

Meiying Gu, Jiawei Zhang, Jiahe Li +4

Recent advances in optimizing Gaussian Splatting for scene geometry have enabled efficient reconstruction of detailed surfaces from images. However, when input views are sparse, su…

cs.CV2024

Prompting Continual Person Search

Pengcheng Zhang, Xiaohan Yu, Xiao Bai +2

The development of person search techniques has been greatly promoted in recent years for its superior practicality and challenging goals. Despite their significant progress, exist…

cs.CV2025

LGD: Leveraging Generative Descriptions for Zero-Shot Referring Image Segmentation

Jiachen Li, Qing Xie, Renshu Gu +3

Zero-shot referring image segmentation aims to locate and segment the target region based on a referring expression, with the primary challenge of aligning and matching semantics a…

cs.CV2022

Feature Fusion Vision Transformer for Fine-Grained Visual Categorization

Jun Wang, Xiaohan Yu, Yongsheng Gao

The core for tackling the fine-grained visual categorization (FGVC) is to learn subtle yet discriminative features. Most previous works achieve this by explicitly selecting the dis…

cs.CV2021

Mask-Guided Feature Extraction and Augmentation for Ultra-Fine-Grained Visual Categorization

Zicheng Pan, Xiaohan Yu, Miaohua Zhang +1

While the fine-grained visual categorization (FGVC) problems have been greatly developed in the past years, the Ultra-fine-grained visual categorization (Ultra-FGVC) problems have…

cs.CV2024

CoR-GS: Sparse-View 3D Gaussian Splatting via Co-Regularization

Jiawei Zhang, Jiahe Li, Xiaohan Yu +4

3D Gaussian Splatting (3DGS) creates a radiance field consisting of 3D Gaussians to represent a scene. With sparse training views, 3DGS easily suffers from overfitting, negatively…

cs.CV2019

Patchy Image Structure Classification Using Multi-Orientation Region Transform

Xiaohan Yu, Yang Zhao, Yongsheng Gao +2

Exterior contour and interior structure are both vital features for classifying objects. However, most of the existing methods consider exterior contour feature and internal struct…

cs.AI2025

AI-SearchPlanner: Modular Agentic Search via Pareto-Optimal Multi-Objective Reinforcement Learning

Lang Mei, Zhihan Yang, Xiaohan Yu +2

Recent studies have explored integrating Large Language Models (LLMs) with search engines to leverage both the LLMs' internal pre-trained knowledge and external information. Specia…

eess.IV2021

EAR-NET: Error Attention Refining Network For Retinal Vessel Segmentation

Jun Wang, Yang Zhao, Linglong Qian +2

The precise detection of blood vessels in retinal images is crucial to the early diagnosis of the retinal vascular diseases, e.g., diabetic, hypertensive and solar retinopathies. E…

cs.CL2025

TableRAG: A Retrieval Augmented Generation Framework for Heterogeneous Document Reasoning

Xiaohan Yu, Pu Jian, Chong Chen

Retrieval-Augmented Generation (RAG) has demonstrated considerable effectiveness in open-domain question answering. However, when applied to heterogeneous documents, comprising bot…

cond-mat.mtrl-sci2025

The Open DAC 2025 Dataset for Sorbent Discovery in Direct Air Capture

Anuroop Sriram, Logan M. Brabson, Xiaohan Yu +12

Identifying useful sorbent materials for direct air capture (DAC) from humid air remains a challenge. We present the Open DAC 2025 (ODAC25) dataset, a significant expansion and imp…

cs.CV2021

Mask Guided Attention For Fine-Grained Patchy Image Classification

Jun Wang, Xiaohan Yu, Yongsheng Gao

In this work, we present a novel mask guided attention (MGA) method for fine-grained patchy image classification. The key challenge of fine-grained patchy image classification lies…

cs.CV2022

PGTRNet: Two-phase Weakly Supervised Object Detection with Pseudo Ground Truth Refinement

Jun Wang, Hefeng Zhou, Xiaohan Yu

Current state-of-the-art weakly supervised object detection (WSOD) studies mainly follow a two-stage training strategy which integrates a fully supervised detector (FSD) with a pur…

cs.CV2025

Improving Medical Visual Representation Learning with Pathological-level Cross-Modal Alignment and Correlation Exploration

Jun Wang, Lixing Zhu, Xiaohan Yu +2

Learning medical visual representations from image-report pairs through joint learning has garnered increasing research attention due to its potential to alleviate the data scarcit…

cs.CV2026

DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents

Huanyao Zhang, Jiepeng Zhou, Runhao Zhao +12

Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametric knowledge limits their ability to address knowledge-intensive…

cs.SE2025

Code Comments for Quantum Software Development Kits: An Empirical Study on Qiskit

Zenghui Zhou, Yuechen Li, Yi Cai +4

Quantum computing is gaining attention from academia and industry. With the quantum Software Development Kits (SDKs), programmers can develop quantum software to explore the power…

cs.ET2025

EDA-Q: Electronic Design Automation for Superconducting Quantum Chip

Bo Zhao, Zhihang Li, Xiaohan Yu +13

Electronic Design Automation (EDA) plays a crucial role in classical chip design and significantly influences the development of quantum chip design. However, traditional EDA tools…

cs.CV2024

EIANet: A Novel Domain Adaptation Approach to Maximize Class Distinction with Neural Collapse Principles

Zicheng Pan, Xiaohan Yu, Yongsheng Gao

Source-free domain adaptation (SFDA) aims to transfer knowledge from a labelled source domain to an unlabelled target domain. A major challenge in SFDA is deriving accurate categor…

cs.IR2025

Break the ID-Language Barrier: An Adaption Framework for LLM-based Sequential Recommendation

Xiaohan Yu, Li Zhang, Xin Zhao +1

The recent breakthrough of large language models (LLMs) in natural language processing has sparked exploration in recommendation systems, however, their limited domain-specific kno…

cs.CL2019

Lattice CNNs for Matching Based Chinese Question Answering

Yuxuan Lai, Yansong Feng, Xiaohan Yu +3

Short text matching often faces the challenges that there are great word mismatch and expression diversity between the two texts, which would be further aggravated in languages lik…

cs.CV2024

Distribution alignment based transfer fusion frameworks on quantum devices for seeking quantum advantages

Xi He, Feiyu Du, Xiaohan Yu +2

The scarcity of labelled data is specifically an urgent challenge in the field of quantum machine learning (QML). Two transfer fusion frameworks are proposed in this paper to predi…

cs.AI2026

BrowseComp-: A Visual, Vertical, and Verifiable Benchmark for Multimodal Browsing Agents

Huanyao Zhang, Jiepeng Zhou, Bo Li +22

Multimodal large language models (MLLMs), equipped with increasingly advanced planning and tool-use capabilities, are evolving into autonomous agents capable of performing multimod…

cs.CV2022

Research on Patch Attentive Neural Process

Xiaohan Yu, Shaochen Mao

Attentive Neural Process (ANP) improves the fitting ability of Neural Process (NP) and improves its prediction accuracy, but the higher time complexity of the model imposes a limit…

cond-mat.mtrl-sci2023

The Open DAC 2023 Dataset and Challenges for Sorbent Discovery in Direct Air Capture

Anuroop Sriram, Sihoon Choi, Xiaohan Yu +6

New methods for carbon dioxide removal are urgently needed to combat global climate change. Direct air capture (DAC) is an emerging technology to capture carbon dioxide directly fr…

cs.CV2019

From Species to Cultivar: Soybean Cultivar Recognition using Multiscale Sliding Chord Matching of Leaf Images

Bin Wang, Yongsheng Gao, Xiaohan Yu +3

Leaf image recognition techniques have been actively researched for plant species identification. However it remains unclear whether leaf patterns can provide sufficient informatio…

cs.CV2025

GeoSVR: Taming Sparse Voxels for Geometrically Accurate Surface Reconstruction

Jiahe Li, Jiawei Zhang, Youmin Zhang +4

Reconstructing accurate surfaces with radiance fields has achieved remarkable progress in recent years. However, prevailing approaches, primarily based on Gaussian Splatting, are i…

cs.CV2026

Latent Anomaly Knowledge Excavation: Unveiling Sparse Sensitive Neurons in Vision-Language Models

Shaotian Li, Shangze Li, Chuancheng Shi +5

Large-scale vision-language models (VLMs) exhibit remarkable zero-shot capabilities, yet the internal mechanisms driving their anomaly detection (AD) performance remain poorly unde…

cs.AI2026

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications

Hao Jiang, Gangtao Xin, Yingdi Huang +35

Large language model agents have advanced rapidly, yet progress remains fragmented across domains, capabilities, task difficulty, and interaction settings. We frame this as full-sc…

cs.IR2024

Multi-Tower Multi-Interest Recommendation with User Representation Repel

Tianyu Xiong, Xiaohan Yu

In the era of information overload, the value of recommender systems has been profoundly recognized in academia and industry alike. Multi-interest sequential recommendation, in par…

cs.RO2025

X2C: A Dataset Featuring Nuanced Facial Expressions for Realistic Humanoid Imitation

Peizhen Li, Longbing Cao, Xiao-Ming Wu +2

The ability to imitate realistic facial expressions is essential for humanoid robots engaged in affective human-robot communication. However, the lack of datasets containing divers…

cs.IR2024

Explainable CTR Prediction via LLM Reasoning

Xiaohan Yu, Li Zhang, Chong Chen

Recommendation Systems have become integral to modern user experiences, but lack transparency in their decision-making processes. Existing explainable recommendation methods are hi…

cs.CV2025

See What You Seek: Semantic Contextual Integration for Cloth-Changing Person Re-Identification

Xiyu Han, Xian Zhong, Wenxin Huang +3

Cloth-changing person re-identification (CC-ReID) aims to match individuals across surveillance cameras despite variations in clothing. Existing methods typically mitigate the impa…

cs.CV2024

Robust Synthetic-to-Real Transfer for Stereo Matching

Jiawei Zhang, Jiahe Li, Lei Huang +4

With advancements in domain generalized stereo matching networks, models pre-trained on synthetic data demonstrate strong robustness to unseen domains. However, few studies have in…

cs.IR2026

SearchArt: Training Long-Horizon Search Agent with Scalable Synthetic and Verified Task

Lang Mei, Xiaohan Yu, Chong Chen +26

Recent advances in large language models (LLMs) have enabled search agents to autonomously tackle complex tasks across extended search and reasoning horizons. However, training eff…

cs.CV2025

Propensity-driven Uncertainty Learning for Sample Exploration in Source-Free Active Domain Adaptation

Zicheng Pan, Xiaohan Yu, Weichuan Zhang +1

Source-free active domain adaptation (SFADA) addresses the challenge of adapting a pre-trained model to new domains without access to source data while minimizing the need for targ…