papers

Publications (50)

cs.CV2019

Learning with Batch-wise Optimal Transport Loss for 3D Shape Recognition

Lin Xu, Han Sun, Yuai Liu

Deep metric learning is essential for visual recognition. The widely used pair-wise (or triplet) based loss objectives cannot make full use of semantical information in training sa…

cs.LG2025

On the Practice of Deep Hierarchical Ensemble Network for Ad Conversion Rate Prediction

Jinfeng Zhuang, Yinrui Li, Runze Su +14

The predictions of click through rate (CTR) and conversion rate (CVR) play a crucial role in the success of ad-recommendation systems. A Deep Hierarchical Ensemble Network (DHEN) h…

cs.RO2025

HANDO: Hierarchical Autonomous Navigation and Dexterous Omni-loco-manipulation

Jingyuan Sun, Chaoran Wang, Mingyu Zhang +6

Seamless loco-manipulation in unstructured environments requires robots to leverage autonomous exploration alongside whole-body control for physical interaction. In this work, we i…

cs.CY2025

Fine-Tuning Large Language Models for Educational Support: Leveraging Gagne's Nine Events of Instruction for Lesson Planning

Linzhao Jia, Changyong Qi, Yuang Wei +2

Effective lesson planning is crucial in education process, serving as the cornerstone for high-quality teaching and the cultivation of a conducive learning atmosphere. This study i…

cs.DB2026

BCTuner: LLM-Guided Monte Carlo Tree Search for Efficient Blockchain Knob Tuning

Yaoyi Deng, Chongyang Tao, Mingxuan Li +4

Knob tuning plays a critical role in improving the performance of permissioned blockchains. However, efficient tuning remains challenging due to the architectural complexity of blo…

cs.CV2022

Attention Guided Network for Salient Object Detection in Optical Remote Sensing Images

Yuhan Lin, Han Sun, Ningzhong Liu +3

Due to the extreme complexity of scale and shape as well as the uncertainty of the predicted location, salient object detection in optical remote sensing images (RSI-SOD) is a very…

cs.LG2024

Continuous Test-time Domain Adaptation for Efficient Fault Detection under Evolving Operating Conditions

Han Sun, Kevin Ammann, Stylianos Giannoulakis +1

Fault detection is crucial in industrial systems to prevent failures and optimize performance by distinguishing abnormal from normal operating conditions. Data-driven methods have…

cs.CR2026

PixJail: Self-Evolving Paper-to-Pipeline Reproduction for Text-to-Image Jailbreak Evaluation

Leyi Sheng, Han Sun, Zhen Sun +4

As Text-to-Image (T2I) jailbreak techniques evolve rapidly, existing benchmarks and reproduction workflows often struggle to keep pace. More importantly, T2I jailbreak evaluation i…

cs.CV2023

SF-FSDA: Source-Free Few-Shot Domain Adaptive Object Detection with Efficient Labeled Data Factory

Han Sun, Rui Gong, Konrad Schindler +1

Domain adaptive object detection aims to leverage the knowledge learned from a labeled source domain to improve the performance on an unlabeled target domain. Prior works typically…

cs.CV2024

B2Net: Camouflaged Object Detection via Boundary Aware and Boundary Fusion

Junmin Cai, Han Sun, Ningzhong Liu

Camouflaged object detection (COD) aims to identify objects in images that are well hidden in the environment due to their high similarity to the background in terms of texture and…

cs.CV2023

SIDE: Self-supervised Intermediate Domain Exploration for Source-free Domain Adaptation

Jiamei Liu, Han Sun, Yizhen Jia +3

Domain adaptation aims to alleviate the domain shift when transferring the knowledge learned from the source domain to the target domain. Due to privacy issues, source-free domain…

cs.LG2026

AFL: A Single-Round Analytic Approach for Federated Learning with Pre-trained Models

Run He, Kai Tong, Di Fang +5

In this paper, we introduce analytic federated learning (AFL), a new training paradigm that brings analytical (i.e., closed-form) solutions to the federated learning (FL) with pre-…

cs.LG2025

The Evolution of Embedding Table Optimization and Multi-Epoch Training in Pinterest Ads Conversion

Andrew Qiu, Shubham Barhate, Hin Wai Lui +7

Deep learning for conversion prediction has found widespread applications in online advertising. These models have become more complex as they are trained to jointly predict multip…

cs.CV2025

Energy-Based Pseudo-Label Refining for Source-free Domain Adaptation

Xinru Meng, Han Sun, Jiamei Liu +2

Source-free domain adaptation (SFDA), which involves adapting models without access to source data, is both demanding and challenging. Existing SFDA techniques typically rely on ps…

cs.CV2025

SLENet: A Guidance-Enhanced Network for Underwater Camouflaged Object Detection

Xinxin Huang, Han Sun, Ningzhong Liu +2

Underwater Camouflaged Object Detection (UCOD) aims to identify objects that blend seamlessly into underwater environments. This task is critically important to marine ecology. How…

cs.LG2025

Entity Representation Learning Through Onsite-Offsite Graph for Pinterest Ads

Jiayin Jin, Zhimeng Pan, Yang Tang +9

Graph Neural Networks (GNN) have been extensively applied to industry recommendation systems, as seen in models like GraphSage\cite{GraphSage}, TwHIM\cite{TwHIM}, LiGNN\cite{LiGNN}…

cs.CV2025

APGNet: Adaptive Prior-Guided for Underwater Camouflaged Object Detection

Xinxin Huang, Han Sun, Junmin Cai +2

Detecting camouflaged objects in underwater environments is crucial for marine ecological research and resource exploration. However, existing methods face two key challenges: unde…

cs.CV2025

SPARNet: Continual Test-Time Adaptation via Sample Partitioning Strategy and Anti-Forgetting Regularization

Xinru Meng, Han Sun, Jiamei Liu +2

Test-time Adaptation (TTA) aims to improve model performance when the model encounters domain changes after deployment. The standard TTA mainly considers the case where the target…

eess.SY2025

Similar Formation Control of Multi-Agent Systems over Directed Acyclic Graphs via Matrix-Weighted Laplacian

Zhipeng Fan, Yujie Xu, Mingyu Fu +3

This brief proposes a distributed formation control strategy via matrix-weighted Laplacian that can achieve a similar formation in 2-D planar using inter-agent relative displacemen…

cs.GR2018

Object-Aware Guidance for Autonomous Scene Reconstruction

Ligang Liu, Xi Xia, Han Sun +5

To carry out autonomous 3D scanning and online reconstruction of unknown indoor scenes, one has to find a balance between global exploration of the entire scene and local scanning…

cs.CV2026

Flow6D: Discrete-to-Continuous Flow Matching for Efficient and Accurate Category-Level 6D Pose Estimation

Mingyu Mei, Li Zhang, Zibo Dai +4

6D pose estimation is a key task in computer vision and embodied AI, widely used in robotic manipulation, augmented reality, etc. Existing methods directly regress in a high-dimens…

cs.CV2021

Robust Ensembling Network for Unsupervised Domain Adaptation

Han Sun, Lei Lin, Ningzhong Liu +1

Recently, in order to address the unsupervised domain adaptation (UDA) problem, extensive studies have been proposed to achieve transferrable models. Among them, the most prevalent…

cs.CL2025

SynDec: A Synthesize-then-Decode Approach for Arbitrary Textual Style Transfer via Large Language Models

Han Sun, Zhen Sun, Zongmin Zhang +3

Large Language Models (LLMs) are emerging as dominant forces for textual style transfer. However, for arbitrary style transfer, LLMs face two key challenges: (1) considerable relia…

cs.CV2021

Multi-scale Edge-based U-shape Network for Salient Object Detection

Han Sun, Yetong Bian, Ningzhong Liu +1

Deep-learning based salient object detection methods achieve great improvements. However, there are still problems existing in the predictions, such as blurry boundary and inaccura…

cs.CR2025

"To Survive, I Must Defect": Jailbreaking LLMs via the Game-Theory Scenarios

Zhen Sun, Zongmin Zhang, Deqi Liang +8

As LLMs become more common, non-expert users can pose risks, prompting extensive research into jailbreak attacks. However, most existing black-box jailbreak attacks rely on hand-cr…

cs.IR2025

Decoupled Entity Representation Learning for Pinterest Ads Ranking

Jie Liu, Yinrui Li, Jiankai Sun +12

In this paper, we introduce a novel framework following an upstream-downstream paradigm to construct user and item (Pin) embeddings from diverse data sources, which are essential f…

cs.IR2026

Fine-Tuned LLM as a Complementary Predictor Improving Ads System

Hui Yang, Daiwei He, Kevin Jiang +20

Recommendation systems power engagement and monetization across feeds, ads, and short-video platforms, but translating the latest advances in Large Language Models into Recommendat…

cs.CV2023

FGFusion: Fine-Grained Lidar-Camera Fusion for 3D Object Detection

Zixuan Yin, Han Sun, Ningzhong Liu +2

Lidars and cameras are critical sensors that provide complementary information for 3D detection in autonomous driving. While most prevalent methods progressively downscale the 3D p…

cs.CV2026

Mitigating Object Hallucinations in LVLMs via Attention Imbalance Rectification

Han Sun, Qin Li, Peixin Wang +1

Object hallucination in Large Vision-Language Models (LVLMs) severely compromises their reliability in real-world applications, posing a critical barrier to their deployment in hig…

cs.CV2025

GeoTexDensifier: Geometry-Texture-Aware Densification for High-Quality Photorealistic 3D Gaussian Splatting

Hanqing Jiang, Xiaojun Xiang, Han Sun +4

3D Gaussian Splatting (3DGS) has recently attracted wide attentions in various areas such as 3D navigation, Virtual Reality (VR) and 3D simulation, due to its photorealistic and ef…

cs.CV2024

DecoratingFusion: A LiDAR-Camera Fusion Network with the Combination of Point-level and Feature-level Fusion

Zixuan Yin, Han Sun, Ningzhong Liu +2

Lidars and cameras play essential roles in autonomous driving, offering complementary information for 3D detection. The state-of-the-art fusion methods integrate them at the featur…

cs.CV2022

Real-Time Elderly Monitoring for Senior Safety by Lightweight Human Action Recognition

Han Sun, Yu Chen

With an increasing number of elders living alone, care-giving from a distance becomes a compelling need, particularly for safety. Real-time monitoring and action recognition are es…

cs.CV2025

DCFS: Continual Test-Time Adaptation via Dual Consistency of Feature and Sample

Wenting Yin, Han Sun, Xinru Meng +2

Continual test-time adaptation aims to continuously adapt a pre-trained model to a stream of target domain data without accessing source data. Without access to source domain data,…

cs.CV2025

FID-Net: A Feature-Enhanced Deep Learning Network for Forest Infestation Detection

Yan Zhang, Baoxin Li, Han Sun +3

Forest pests threaten ecosystem stability, requiring efficient monitoring. To overcome the limitations of traditional methods in large-scale, fine-grained detection, this study foc…

cs.LG2025

From Physics to Machine Learning and Back: Part II - Learning and Observational Bias in PHM

Olga Fink, Ismail Nejjar, Vinay Sharma +13

Prognostics and Health Management ensures the reliability, safety, and efficiency of complex engineered systems by enabling fault detection, anticipating equipment failures, and op…

physics.soc-ph2007

A limited resource model of fault-tolerant capability against cascading failure of complex network

Ping Li, Bing-Hong Wang, Han Sun +2

We propose a novel capacity model for complex networks against cascading failure. In this model, vertices with both higher loads and larger degrees should be paid more extra capaci…

cs.CV2025

SliceSemOcc: Vertical Slice Based Multimodal 3D Semantic Occupancy Representation

Han Huang, Han Sun, Ningzhong Liu +2

Driven by autonomous driving's demands for precise 3D perception, 3D semantic occupancy prediction has become a pivotal research topic. Unlike bird's-eye-view (BEV) methods, which…

cs.RO2026

Exploring Pose-Guided Imitation Learning for Robotic Precise Insertion

Han Sun, Sheng Liu, Yizhao Wang +5

Imitation learning is promising for robotic manipulation, but \emph{precise insertion} in the real world remains difficult due to contact-rich dynamics, tight clearances, and limit…

stat.AP2025

TARD: Test-time Domain Adaptation for Robust Fault Detection under Evolving Operating Conditions

Han Sun, Olga Fink

Fault detection is essential in complex industrial systems to prevent failures and optimize performance by distinguishing abnormal from normal operating conditions. With the growin…

cs.CV2023

SimMMDG: A Simple and Effective Framework for Multi-modal Domain Generalization

Hao Dong, Ismail Nejjar, Han Sun +2

In real-world scenarios, achieving domain generalization (DG) presents significant challenges as models are required to generalize to unknown target distributions. Generalizing to…

cs.CV2023

Prompting Diffusion Representations for Cross-Domain Semantic Segmentation

Rui Gong, Martin Danelljan, Han Sun +2

While originally designed for image generation, diffusion models have recently shown to provide excellent pretrained feature representations for semantic segmentation. Intrigued by…

cs.CV2022

MonoSIM: Simulating Learning Behaviors of Heterogeneous Point Cloud Object Detectors for Monocular 3D Object Detection

Han Sun, Zhaoxin Fan, Zhenbo Song +3

Monocular 3D object detection is a fundamental but very important task to many applications including autonomous driving, robotic grasping and augmented reality. Existing leading m…

cs.RO2026

ActivePose: Active 6D Object Pose Estimation and Tracking for Robotic Manipulation

Sheng Liu, Zhe Li, Weiheng Wang +6

Accurate 6-DoF object pose estimation and tracking are critical for reliable robotic manipulation. However, zero-shot methods often fail under viewpoint-induced ambiguities and fix…

cs.CV2021

MPI: Multi-receptive and Parallel Integration for Salient Object Detection

Han Sun, Jun Cen, Ningzhong Liu +2

The semantic representation of deep features is essential for image context understanding, and effective fusion of features with different semantic representations can significantl…

cs.CV2022

Polycentric Clustering and Structural Regularization for Source-free Unsupervised Domain Adaptation

Xinyu Guan, Han Sun, Ningzhong Liu +1

Source-Free Domain Adaptation (SFDA) aims to solve the domain adaptation problem by transferring the knowledge learned from a pre-trained source model to an unseen target domain. M…

cs.CV2025

Unseen Visual Anomaly Generation

Han Sun, Yunkang Cao, Hao Dong +1

Visual anomaly detection (AD) presents significant challenges due to the scarcity of anomalous data samples. While numerous works have been proposed to synthesize anomalous samples…

cs.IR2025

Multi-Faceted Large Embedding Tables for Pinterest Ads Ranking

Runze Su, Jiayin Jin, Jiacheng Li +16

Large embedding tables are indispensable in modern recommendation systems, thanks to their ability to effectively capture and memorize intricate details of interactions among diver…

cs.CV2022

A Task-aware Dual Similarity Network for Fine-grained Few-shot Learning

Yan Qi, Han Sun, Ningzhong Liu +1

The goal of fine-grained few-shot learning is to recognize sub-categories under the same super-category by learning few labeled samples. Most of the recent approaches adopt a singl…

cs.CV2025

DynAlign: Unsupervised Dynamic Taxonomy Alignment for Cross-Domain Segmentation

Han Sun, Rui Gong, Ismail Nejjar +1

Current unsupervised domain adaptation (UDA) methods for semantic segmentation typically assume identical class labels between the source and target domains. This assumption ignore…

cs.CV2022

A lightweight multi-scale context network for salient object detection in optical remote sensing images

Yuhan Lin, Han Sun, Ningzhong Liu +3

Due to the more dramatic multi-scale variations and more complicated foregrounds and backgrounds in optical remote sensing images (RSIs), the salient object detection (SOD) for opt…