papers

Publications (23)

cs.CV2020

Classification Calibration for Long-tail Instance Segmentation

Tao Wang, Yu Li, Bingyi Kang +5

Remarkable progress has been made in object instance detection and segmentation in recent years. However, existing state-of-the-art methods are mostly evaluated with fairly balance…

cs.CV2019

CGNet: A Light-weight Context Guided Network for Semantic Segmentation

Tianyi Wu, Sheng Tang, Rui Zhang +1

The demand of applying semantic segmentation model on mobile devices has been increasing rapidly. Current state-of-the-art networks have enormous amount of parameters hence unsuita…

cs.CV2020

Overcoming Classifier Imbalance for Long-tail Object Detection with Balanced Group Softmax

Yu Li, Tao Wang, Bingyi Kang +4

Solving long-tail large vocabulary object detection with deep learning based models is a challenging and demanding task, which is however under-explored.In this work, we provide th…

cs.CV2025

Beyond Semantic Features: Pixel-level Mapping for Generalized AI-Generated Image Detection

Chenming Zhou, Jiaan Wang, Yu Li +3

The rapid evolution of generative technologies necessitates reliable methods for detecting AI-generated images. A critical limitation of current detectors is their failure to gener…

astro-ph.HE2026

GeV γ-ray emission in the low-mass star-forming region AFGL 490

Li-Nuo Yang, Sheng Tang, Pak-Hin Thomas Tam

We report the discovery of an extended GeV γ-ray source, 4FGL J0330.7+5845e, associated with the star-forming region AFGL 490 using 17 years of Fermi-LAT data. The emission is spa…

cs.LG2025

Beyond Token-level Supervision: Unlocking the Potential of Decoding-based Regression via Reinforcement Learning

Ming Chen, Sheng Tang, Rong-Xi Tan +4

Decoding-based regression, which reformulates regression as a sequence generation task, has emerged as a promising paradigm of applying large language models for numerical predicti…

cs.LG2025

Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models

Ren-Jian Wang, Ke Xue, Zeyu Qin +7

Ensuring the safety and robustness of large language models (LLMs) is a fundamental challenge and a critical prerequisite for the responsible deployment of artificial intelligence.…

cs.CV2026

TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy

Hao Sun, Hao Yan, Mengting Chen +7

While Video Virtual Try-on (VVT) has achieved remarkable progress in synthesizing realistic garment overlays on dynamic subjects, existing paradigms remains fundamentally constrain…

cs.CV2024

Learning Monocular Depth from Events via Egomotion Compensation

Haitao Meng, Chonghao Zhong, Sheng Tang +6

Event cameras are neuromorphically inspired sensors that sparsely and asynchronously report brightness changes. Their unique characteristics of high temporal resolution, high dynam…

cs.CV2020

The Devil is in Classification: A Simple Framework for Long-tail Object Detection and Instance Segmentation

Tao Wang, Yu Li, Bingyi Kang +5

Most existing object instance detection and segmentation models only work well on fairly balanced benchmarks where per-category training sample numbers are comparable, such as COCO…

cs.IR2015

HDIdx: High-Dimensional Indexing for Efficient Approximate Nearest Neighbor Search

Ji Wan, Sheng Tang, Yongdong Zhang +3

Fast Nearest Neighbor (NN) search is a fundamental challenge in large-scale data processing and analytics, particularly for analyzing multimedia contents which are often of high di…

cs.CV2025

Attend to Not Attended: Structure-then-Detail Token Merging for Post-training DiT Acceleration

Haipeng Fang, Sheng Tang, Juan Cao +3

Diffusion transformers have shown exceptional performance in visual generation but incur high computational costs. Token reduction techniques that compress models by sharing the de…

cs.CV2020

Visual Relation Grounding in Videos

Junbin Xiao, Xindi Shang, Xun Yang +2

In this paper, we explore a novel task named visual Relation Grounding in Videos (vRGV). The task aims at spatio-temporally localizing the given relations in the form of subject-pr…

cs.CV2024

DragEntity: Trajectory Guided Video Generation using Entity and Positional Relationships

Zhang Wan, Sheng Tang, Jiawei Wei +2

In recent years, diffusion models have achieved tremendous success in the field of video generation, with controllable video generation receiving significant attention. However, ex…

cs.CV2026

Fleet: Few Shots Lead Effective AI-generated Image Detection

Jiaan Wang, Sirui Liu, Yu Li +3

AI-generated image (AIGI) detection is undergoing a critical transition from laboratory benchmarks to open-world adversarial defense. The prevalent paradigm focuses on finding stat…

cs.CV2024

Topology-preserving Adversarial Training for Alleviating Natural Accuracy Degradation

Xiaoyue Mi, Fan Tang, Yepeng Weng +5

Despite the effectiveness in improving the robustness of neural networks, adversarial training has suffered from the natural accuracy degradation problem, i.e., accuracy on natural…

cs.CV2018

Tree-structured Kronecker Convolutional Network for Semantic Segmentation

Tianyi Wu, Sheng Tang, Rui Zhang +2

Most existing semantic segmentation methods employ atrous convolution to enlarge the receptive field of filters, but neglect partial information. To tackle this issue, we firstly p…

cs.CV2021

Learning to Disentangle GAN Fingerprint for Fake Image Attribution

Tianyun Yang, Juan Cao, Qiang Sheng +4

Rapid pace of generative models has brought about new threats to visual forensics such as malicious personation and digital copyright infringement, which promotes works on fake ima…

cs.CV2023

Dance Your Latents: Consistent Dance Generation through Spatial-temporal Subspace Attention Guided by Motion Flow

Haipeng Fang, Zhihao Sun, Ziyao Huang +3

The advancement of generative AI has extended to the realm of Human Dance Generation, demonstrating superior generative capacities. However, current methods still exhibit deficienc…

cs.CV2019

Asymmetric GAN for Unpaired Image-to-image Translation

Yu Li, Sheng Tang, Rui Zhang +3

Unpaired image-to-image translation problem aims to model the mapping from one domain to another with unpaired training data. Current works like the well-acknowledged Cycle GAN pro…

cs.CV2018

Style Separation and Synthesis via Generative Adversarial Networks

Rui Zhang, Sheng Tang, Yu Li +4

Style synthesis attracts great interests recently, while few works focus on its dual problem "style separation". In this paper, we propose the Style Separation and Synthesis Genera…

cs.CV2023

Progressive Open Space Expansion for Open-Set Model Attribution

Tianyun Yang, Danding Wang, Fan Tang +3

Despite the remarkable progress in generative technology, the Janus-faced issues of intellectual property protection and malicious content supervision have arisen. Efforts have bee…

cs.CV2019

Consensus Feature Network for Scene Parsing

Tianyi Wu, Sheng Tang, Rui Zhang +2

Scene parsing is challenging as it aims to assign one of the semantic categories to each pixel in scene images. Thus, pixel-level features are desired for scene parsing. However, c…