papers

Publications (31)

cs.CV2017

Integrated Deep and Shallow Networks for Salient Object Detection

Jing Zhang, Bo Li, Yuchao Dai +2

Deep convolutional neural network (CNN) based salient object detection methods have achieved state-of-the-art performance and outperform those unsupervised methods with a wide marg…

cs.CV2020

Self-supervised Modal and View Invariant Feature Learning

Longlong Jing, Yucheng Chen, Ling Zhang +2

Most of the existing self-supervised feature learning methods for 3D data either learn 3D features from point cloud data or from multi-view images. By exploring the inherent multi-…

cs.CV2026

Envisioning global urban development with satellite imagery and generative AI

Kailai Sun, Yuebing Liang, Mingyi He +5

Urban development has been a defining force in human history, shaping cities for centuries. However, past studies mostly analyze such development as predictive tasks, failing to re…

cs.CV2026

SlideCheck: Guiding Self-Supervised Pretraining of Pathology Foundation Models via Dataset Distributions

Mingyi He, Xinyi Guo, Xitong Ling +7

Pathology foundation models are pretrained on large streams of WSI-derived patches, while supervision during data construction is often slide-level, sparse, or heterogeneous. This…

cs.CV2017

Skeleton Boxes: Solving skeleton based action detection with a single deep convolutional neural network

Bo Li, Huahui Chen, Yucheng Chen +2

Action recognition from well-segmented 3D skeleton video has been intensively studied. However, due to the difficulty in representing the 3D skeleton video and the lack of training…

cs.CV2017

Monocular Depth Estimation with Hierarchical Fusion of Dilated CNNs and Soft-Weighted-Sum Inference

Bo Li, Yuchao Dai, Mingyi He

Monocular depth estimation is a challenging task in complex compositions depicting multiple objects of diverse scales. Albeit the recent great progress thanks to the deep convoluti…

cs.CV2022

A Representation Separation Perspective to Correspondences-free Unsupervised 3D Point Cloud Registration

Zhiyuan Zhang, Jiadai Sun, Yuchao Dai +3

3D point cloud registration in remote sensing field has been greatly advanced by deep learning based methods, where the rigid transformation is either directly regressed from the t…

cs.CV2017

Single image depth estimation by dilated deep residual convolutional neural network and soft-weight-sum inference

Bo Li, Yuchao Dai, Huahui Chen +1

This paper proposes a new residual convolutional neural network (CNN) architecture for single image depth estimation. Compared with existing deep CNN based methods, our method achi…

cs.AI2025

Generative AI for Urban Design: A Stepwise Approach Integrating Human Expertise with Multimodal Diffusion Models

Mingyi He, Yuebing Liang, Shenhao Wang +5

Urban design is a multifaceted process that demands careful consideration of site-specific constraints and collaboration among diverse professionals and stakeholders. The advent of…

cs.CV2022

End-to-end Learning the Partial Permutation Matrix for Robust 3D Point Cloud Registration

Zhiyuan Zhang, Jiadai Sun, Yuchao Dai +3

Even though considerable progress has been made in deep learning-based 3D point cloud processing, how to obtain accurate correspondences for robust registration remains a major cha…

cs.CV2021

SUNet: Symmetric Undistortion Network for Rolling Shutter Correction

Bin Fan, Yuchao Dai, Mingyi He

The vast majority of modern consumer-grade cameras employ a rolling shutter mechanism, leading to image distortions if the camera moves during image acquisition. In this paper, we…

cs.CV2022

Context-Aware Video Reconstruction for Rolling Shutter Cameras

Bin Fan, Yuchao Dai, Zhiyuan Zhang +2

With the ubiquity of rolling shutter (RS) cameras, it is becoming increasingly attractive to recover the latent global shutter (GS) video from two consecutive RS frames, which also…

cs.CV2017

Dense Non-rigid Structure-from-Motion Made Easy - A Spatial-Temporal Smoothness based Solution

Yuchao Dai, Huizhong Deng, Mingyi He

This paper proposes a simple spatial-temporal smoothness based method for solving dense non-rigid structure-from-motion (NRSfM). First, we revisit the temporal smoothness and demon…

physics.soc-ph2019

Pattern and Anomaly Detection in Urban Temporal Networks

Mingyi He, Shivam Pathak, Urwa Muaz +4

Broad spectrum of urban activities including mobility can be modeled as temporal networks evolving over time. Abrupt changes in urban dynamics caused by events such as disruption o…

cs.CV2019

MSDC-Net: Multi-Scale Dense and Contextual Networks for Automated Disparity Map for Stereo Matching

Zhibo Rao, Mingyi He, Yuchao Dai +3

Disparity prediction from stereo images is essential to computer vision applications including autonomous driving, 3D model reconstruction, and object detection. To predict accurat…

cs.CV2017

Skeleton based action recognition using translation-scale invariant image mapping and multi-scale deep cnn

Bo Li, Mingyi He, Xuelian Cheng +2

This paper presents an image classification based approach for skeleton-based video action recognition problem. Firstly, A dataset independent translation-scale invariant image map…

cs.CV2019

Multi-scale Cross-form Pyramid Network for Stereo Matching

Zhidong Zhu, Mingyi He, Yuchao Dai +2

Stereo matching plays an indispensable part in autonomous driving, robotics and 3D scene reconstruction. We propose a novel deep learning architecture, which called CFP-Net, a Cros…

physics.soc-ph2024

Cities Reconceptualized: Unveiling Hidden Uniform Urban Shape through Commute Flow Modeling in Major US Cities

Margarita Mishina, Mingyi He, Venu Garikapati +1

Urban development is shaped by historical, geographical, and economic factors, presenting challenges for planners in understanding urban form. This study models commute flows acros…

cs.CV2023

Transferable Attack for Semantic Segmentation

Mengqi He, Jing Zhang, Zhaoyuan Yang +3

We analysis performance of semantic segmentation models wrt. adversarial attacks, and observe that the adversarial examples generated from a source model fail to attack the target…

cs.CV2026

Is Class Signal Clustered or Routed in Task-Induced Implicit Neural Representation Weight Spaces?

Xinyi Guo, Mingyi He, Haobin Ding +7

Implicit neural representations (INRs) encode images as neural-network weights, making image classification a problem of weight-space classifiability. A natural geometric hypothesi…

cs.CV2022

VRNet: Learning the Rectified Virtual Corresponding Points for 3D Point Cloud Registration

Zhiyuan Zhang, Jiadai Sun, Yuchao Dai +2

3D point cloud registration is fragile to outliers, which are labeled as the points without corresponding points. To handle this problem, a widely adopted strategy is to estimate t…

cs.LG2021

Dense Uncertainty Estimation

Jing Zhang, Yuchao Dai, Mochu Xiang +7

Deep neural networks can be roughly divided into deterministic neural networks and stochastic neural networks.The former is usually trained to achieve a mapping from input space to…

cs.CV2020

Monocular Human Pose Estimation: A Survey of Deep Learning-based Methods

Yucheng Chen, Yingli Tian, Mingyi He

Vision-based monocular human pose estimation, as one of the most fundamental and challenging problems in computer vision, aims to obtain posture of the human body from input images…

cs.CV2024

LoFLAT: Local Feature Matching using Focused Linear Attention Transformer

Naijian Cao, Renjie He, Yuchao Dai +1

Local feature matching is an essential technique in image matching and plays a critical role in a wide range of vision-based applications. However, existing Transformer-based detec…

cs.CV2024

A Revisit of the Normalized Eight-Point Algorithm and A Self-Supervised Deep Solution

Bin Fan, Yuchao Dai, Yongduek Seo +1

The normalized eight-point algorithm has been widely viewed as the cornerstone in two-view geometry computation, where the seminal Hartley's normalization has greatly improved the…

cs.CV2020

Self-supervised Feature Learning by Cross-modality and Cross-view Correspondences

Longlong Jing, Yucheng Chen, Ling Zhang +2

The success of supervised learning requires large-scale ground truth labels which are very expensive, time-consuming, or may need special skills to annotate. To address this issue,…

cs.CV2017

Deep Edge-Aware Saliency Detection

Jing Zhang, Yuchao Dai, Fatih Porikli +1

There has been profound progress in visual saliency thanks to the deep learning architectures, however, there still exist three major challenges that hinder the detection performan…

cs.CV2026

SENSE: Satellite-based ENergy Synthesis for Sustainable Environment

Kailai Sun, Mingyi He, Heye Huang +5

Urban Building Energy Modeling plays a critical role in achieving the United Nations' Sustainable Development Goals 7 and 11. Although existing studies based on satellite imagery a…

physics.data-an2021

Pattern Ensembling for Spatial Trajectory Reconstruction

Shivam Pathak, Mingyi He, Sergey Malinchik +1

Digital sensing provides an unprecedented opportunity to assess and understand mobility. However, incompleteness, missing information, possible inaccuracies, and temporal heterogen…

cs.CV2023

Joint Salient Object Detection and Camouflaged Object Detection via Uncertainty-aware Learning

Aixuan Li, Jing Zhang, Yunqiu Lv +4

Salient objects attract human attention and usually stand out clearly from their surroundings. In contrast, camouflaged objects share similar colors or textures with the environmen…

cs.CV2022

Learning a Task-specific Descriptor for Robust Matching of 3D Point Clouds

Zhiyuan Zhang, Yuchao Dai, Bin Fan +2

Existing learning-based point feature descriptors are usually task-agnostic, which pursue describing the individual 3D point clouds as accurate as possible. However, the matching t…