Publications (44)
A Compositional Feature Embedding and Similarity Metric for Ultra-Fine-Grained Visual Categorization
Yajie Sun, Miaohua Zhang, Xiaohan Yu +2
Fine-grained visual categorization (FGVC), which aims at classifying objects with small inter-class variances, has been significantly advanced in recent years. However, ultra-fine-…
MSearcher: Modular Multimodal Information Seeking Agency with Retrieval-Oriented Reasoning
Xiaohan Yu, Chao Feng, Lang Mei +1
Recent advances in DeepResearch-style agents have demonstrated strong capabilities in autonomous information acquisition and synthesize from real-world web environments. However, e…
CogPlanner: Unveiling the Potential of Agentic Multimodal Retrieval Augmented Generation with Planning
Xiaohan Yu, Zhihan Yang, Chong Chen
Multimodal Retrieval Augmented Generation (MRAG) systems have shown promise in enhancing the generation capabilities of multimodal large language models (MLLMs). However, existing…
Revisiting Photometric Ambiguity for Accurate Gaussian-Splatting Surface Reconstruction
Jiahe Li, Jiawei Zhang, Xiao Bai +4
Surface reconstruction with differentiable rendering has achieved impressive performance in recent years, yet the pervasive photometric ambiguities have strictly bottlenecked exist…
SCE-SLAM: Scale-Consistent Monocular SLAM via Scene Coordinate Embeddings
Yuchen Wu, Jiahe Li, Xiaohan Yu +3
Monocular visual SLAM enables 3D reconstruction from internet video and autonomous navigation on resource-constrained platforms, yet suffers from scale drift, i.e., the gradual div…
RA-Rec: An Efficient ID Representation Alignment Framework for LLM-based Recommendation
Xiaohan Yu, Li Zhang, Xin Zhao +2
Large language models (LLM) have recently emerged as a powerful tool for a variety of natural language processing tasks, bringing a new surge of combining LLM with recommendation s…
UGotMe: An Embodied System for Affective Human-Robot Interaction
Peizhen Li, Longbing Cao, Xiao-Ming Wu +2
Equipping humanoid robots with the capability to understand emotional states of human interactants and express emotions appropriately according to situations is essential for affec…
SpectralKAN: Weighted Activation Distribution Kolmogorov-Arnold Network for Hyperspectral Image Change Detection
Yanheng Wang, Xiaohan Yu, Yongsheng Gao +6
Kolmogorov-Arnold networks (KANs) represent data features by learning the activation functions and demonstrate superior accuracy with fewer parameters, FLOPs, GPU memory usage (Mem…
Visual and Semantic Prompt Collaboration for Generalized Zero-Shot Learning
Huajie Jiang, Zhengxian Li, Xiaohan Yu +4
Generalized zero-shot learning aims to recognize both seen and unseen classes with the help of semantic information that is shared among different classes. It inevitably requires c…
SparseSurf: Sparse-View 3D Gaussian Splatting for Surface Reconstruction
Meiying Gu, Jiawei Zhang, Jiahe Li +4
Recent advances in optimizing Gaussian Splatting for scene geometry have enabled efficient reconstruction of detailed surfaces from images. However, when input views are sparse, su…
Prompting Continual Person Search
Pengcheng Zhang, Xiaohan Yu, Xiao Bai +2
The development of person search techniques has been greatly promoted in recent years for its superior practicality and challenging goals. Despite their significant progress, exist…
LGD: Leveraging Generative Descriptions for Zero-Shot Referring Image Segmentation
Jiachen Li, Qing Xie, Renshu Gu +3
Zero-shot referring image segmentation aims to locate and segment the target region based on a referring expression, with the primary challenge of aligning and matching semantics a…
Feature Fusion Vision Transformer for Fine-Grained Visual Categorization
Jun Wang, Xiaohan Yu, Yongsheng Gao
The core for tackling the fine-grained visual categorization (FGVC) is to learn subtle yet discriminative features. Most previous works achieve this by explicitly selecting the dis…
Mask-Guided Feature Extraction and Augmentation for Ultra-Fine-Grained Visual Categorization
Zicheng Pan, Xiaohan Yu, Miaohua Zhang +1
While the fine-grained visual categorization (FGVC) problems have been greatly developed in the past years, the Ultra-fine-grained visual categorization (Ultra-FGVC) problems have…
CoR-GS: Sparse-View 3D Gaussian Splatting via Co-Regularization
Jiawei Zhang, Jiahe Li, Xiaohan Yu +4
3D Gaussian Splatting (3DGS) creates a radiance field consisting of 3D Gaussians to represent a scene. With sparse training views, 3DGS easily suffers from overfitting, negatively…
Patchy Image Structure Classification Using Multi-Orientation Region Transform
Xiaohan Yu, Yang Zhao, Yongsheng Gao +2
Exterior contour and interior structure are both vital features for classifying objects. However, most of the existing methods consider exterior contour feature and internal struct…
AI-SearchPlanner: Modular Agentic Search via Pareto-Optimal Multi-Objective Reinforcement Learning
Lang Mei, Zhihan Yang, Xiaohan Yu +2
Recent studies have explored integrating Large Language Models (LLMs) with search engines to leverage both the LLMs' internal pre-trained knowledge and external information. Specia…
EAR-NET: Error Attention Refining Network For Retinal Vessel Segmentation
Jun Wang, Yang Zhao, Linglong Qian +2
The precise detection of blood vessels in retinal images is crucial to the early diagnosis of the retinal vascular diseases, e.g., diabetic, hypertensive and solar retinopathies. E…
TableRAG: A Retrieval Augmented Generation Framework for Heterogeneous Document Reasoning
Xiaohan Yu, Pu Jian, Chong Chen
Retrieval-Augmented Generation (RAG) has demonstrated considerable effectiveness in open-domain question answering. However, when applied to heterogeneous documents, comprising bot…
The Open DAC 2025 Dataset for Sorbent Discovery in Direct Air Capture
Anuroop Sriram, Logan M. Brabson, Xiaohan Yu +12
Identifying useful sorbent materials for direct air capture (DAC) from humid air remains a challenge. We present the Open DAC 2025 (ODAC25) dataset, a significant expansion and imp…
Mask Guided Attention For Fine-Grained Patchy Image Classification
Jun Wang, Xiaohan Yu, Yongsheng Gao
In this work, we present a novel mask guided attention (MGA) method for fine-grained patchy image classification. The key challenge of fine-grained patchy image classification lies…
PGTRNet: Two-phase Weakly Supervised Object Detection with Pseudo Ground Truth Refinement
Jun Wang, Hefeng Zhou, Xiaohan Yu
Current state-of-the-art weakly supervised object detection (WSOD) studies mainly follow a two-stage training strategy which integrates a fully supervised detector (FSD) with a pur…
Improving Medical Visual Representation Learning with Pathological-level Cross-Modal Alignment and Correlation Exploration
Jun Wang, Lixing Zhu, Xiaohan Yu +2
Learning medical visual representations from image-report pairs through joint learning has garnered increasing research attention due to its potential to alleviate the data scarcit…
DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents
Huanyao Zhang, Jiepeng Zhou, Runhao Zhao +12
Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametric knowledge limits their ability to address knowledge-intensive…
Code Comments for Quantum Software Development Kits: An Empirical Study on Qiskit
Zenghui Zhou, Yuechen Li, Yi Cai +4
Quantum computing is gaining attention from academia and industry. With the quantum Software Development Kits (SDKs), programmers can develop quantum software to explore the power…
EDA-Q: Electronic Design Automation for Superconducting Quantum Chip
Bo Zhao, Zhihang Li, Xiaohan Yu +13
Electronic Design Automation (EDA) plays a crucial role in classical chip design and significantly influences the development of quantum chip design. However, traditional EDA tools…
EIANet: A Novel Domain Adaptation Approach to Maximize Class Distinction with Neural Collapse Principles
Zicheng Pan, Xiaohan Yu, Yongsheng Gao
Source-free domain adaptation (SFDA) aims to transfer knowledge from a labelled source domain to an unlabelled target domain. A major challenge in SFDA is deriving accurate categor…
Break the ID-Language Barrier: An Adaption Framework for LLM-based Sequential Recommendation
Xiaohan Yu, Li Zhang, Xin Zhao +1
The recent breakthrough of large language models (LLMs) in natural language processing has sparked exploration in recommendation systems, however, their limited domain-specific kno…
Lattice CNNs for Matching Based Chinese Question Answering
Yuxuan Lai, Yansong Feng, Xiaohan Yu +3
Short text matching often faces the challenges that there are great word mismatch and expression diversity between the two texts, which would be further aggravated in languages lik…
Distribution alignment based transfer fusion frameworks on quantum devices for seeking quantum advantages
Xi He, Feiyu Du, Xiaohan Yu +2
The scarcity of labelled data is specifically an urgent challenge in the field of quantum machine learning (QML). Two transfer fusion frameworks are proposed in this paper to predi…
BrowseComp-: A Visual, Vertical, and Verifiable Benchmark for Multimodal Browsing Agents
Huanyao Zhang, Jiepeng Zhou, Bo Li +22
Multimodal large language models (MLLMs), equipped with increasingly advanced planning and tool-use capabilities, are evolving into autonomous agents capable of performing multimod…
Research on Patch Attentive Neural Process
Xiaohan Yu, Shaochen Mao
Attentive Neural Process (ANP) improves the fitting ability of Neural Process (NP) and improves its prediction accuracy, but the higher time complexity of the model imposes a limit…
The Open DAC 2023 Dataset and Challenges for Sorbent Discovery in Direct Air Capture
Anuroop Sriram, Sihoon Choi, Xiaohan Yu +6
New methods for carbon dioxide removal are urgently needed to combat global climate change. Direct air capture (DAC) is an emerging technology to capture carbon dioxide directly fr…
From Species to Cultivar: Soybean Cultivar Recognition using Multiscale Sliding Chord Matching of Leaf Images
Bin Wang, Yongsheng Gao, Xiaohan Yu +3
Leaf image recognition techniques have been actively researched for plant species identification. However it remains unclear whether leaf patterns can provide sufficient informatio…
GeoSVR: Taming Sparse Voxels for Geometrically Accurate Surface Reconstruction
Jiahe Li, Jiawei Zhang, Youmin Zhang +4
Reconstructing accurate surfaces with radiance fields has achieved remarkable progress in recent years. However, prevailing approaches, primarily based on Gaussian Splatting, are i…
Latent Anomaly Knowledge Excavation: Unveiling Sparse Sensitive Neurons in Vision-Language Models
Shaotian Li, Shangze Li, Chuancheng Shi +5
Large-scale vision-language models (VLMs) exhibit remarkable zero-shot capabilities, yet the internal mechanisms driving their anomaly detection (AD) performance remain poorly unde…
AgentOmnia: Scaling Agentic Models for Full-Scenario Applications
Hao Jiang, Gangtao Xin, Yingdi Huang +35
Large language model agents have advanced rapidly, yet progress remains fragmented across domains, capabilities, task difficulty, and interaction settings. We frame this as full-sc…
Multi-Tower Multi-Interest Recommendation with User Representation Repel
Tianyu Xiong, Xiaohan Yu
In the era of information overload, the value of recommender systems has been profoundly recognized in academia and industry alike. Multi-interest sequential recommendation, in par…
X2C: A Dataset Featuring Nuanced Facial Expressions for Realistic Humanoid Imitation
Peizhen Li, Longbing Cao, Xiao-Ming Wu +2
The ability to imitate realistic facial expressions is essential for humanoid robots engaged in affective human-robot communication. However, the lack of datasets containing divers…
Explainable CTR Prediction via LLM Reasoning
Xiaohan Yu, Li Zhang, Chong Chen
Recommendation Systems have become integral to modern user experiences, but lack transparency in their decision-making processes. Existing explainable recommendation methods are hi…
See What You Seek: Semantic Contextual Integration for Cloth-Changing Person Re-Identification
Xiyu Han, Xian Zhong, Wenxin Huang +3
Cloth-changing person re-identification (CC-ReID) aims to match individuals across surveillance cameras despite variations in clothing. Existing methods typically mitigate the impa…
Robust Synthetic-to-Real Transfer for Stereo Matching
Jiawei Zhang, Jiahe Li, Lei Huang +4
With advancements in domain generalized stereo matching networks, models pre-trained on synthetic data demonstrate strong robustness to unseen domains. However, few studies have in…
SearchArt: Training Long-Horizon Search Agent with Scalable Synthetic and Verified Task
Lang Mei, Xiaohan Yu, Chong Chen +26
Recent advances in large language models (LLMs) have enabled search agents to autonomously tackle complex tasks across extended search and reasoning horizons. However, training eff…
Propensity-driven Uncertainty Learning for Sample Exploration in Source-Free Active Domain Adaptation
Zicheng Pan, Xiaohan Yu, Weichuan Zhang +1
Source-free active domain adaptation (SFADA) addresses the challenge of adapting a pre-trained model to new domains without access to source data while minimizing the need for targ…