papers

Publications (67)

cs.RO2026

Robotic Manipulation is Vision-to-Geometry Mapping (): Vision-Geometry Backbones over Language and Video Models

Zijian Song, Qichang Li, Jiawei Zhou +4

At its core, robotic manipulation is a problem of vision-to-geometry mapping (). Physical actions are fundamentally defined by geometric properties like 3D posi…

cs.CV2018

Knowledge-Embedded Representation Learning for Fine-Grained Image Recognition

Tianshui Chen, Liang Lin, Riquan Chen +2

Humans can naturally understand an image in depth with the aid of rich knowledge accumulated from daily lives or professions. For example, to achieve fine-grained image recognition…

cs.CV2024

Globally Correlation-Aware Hard Negative Generation

Wenjie Peng, Hongxiang Huang, Tianshui Chen +3

Hard negative generation aims to generate informative negative samples that help to determine the decision boundaries and thus facilitate advancing deep metric learning. Current wo…

cs.CV2026

Personalized Cross-Modal Emotional Correlation Learning for Speech-Preserving Facial Expression Manipulation

Tianshui Chen, Yujie Zhu, Jianman Lin +4

Speech-preserving facial expression manipulation (SPFEM) aims to enhance human expressiveness without altering mouth movements tied to the original speech. A primary challenge in t…

cs.CV2020

Knowledge Graph Transfer Network for Few-Shot Recognition

Riquan Chen, Tianshui Chen, Xiaolu Hui +3

Few-shot learning aims to learn novel categories from very few samples given some base categories with sufficient training samples. The main challenge of this task is the novel cat…

cs.CV2023

OccluMix: Towards De-Occlusion Virtual Try-on by Semantically-Guided Mixup

Zhijing Yang, Junyang Chen, Yukai Shi +3

Image Virtual try-on aims at replacing the cloth on a personal image with a garment image (in-shop clothes), which has attracted increasing attention from the multimedia and comput…

cs.CV2017

Multi-label Image Recognition by Recurrently Discovering Attentional Regions

Zhouxia Wang, Tianshui Chen, Guanbin Li +2

This paper proposes a novel deep architecture to address multi-label image recognition, a fundamental and practical task towards general visual understanding. Current solutions for…

cs.CV2022

Semantic-Aware Representation Blending for Multi-Label Image Recognition with Partial Labels

Tao Pu, Tianshui Chen, Hefeng Wu +1

Training the multi-label image recognition models with partial labels, in which merely some labels are known while others are unknown for each image, is a considerably challenging…

cs.CV2024

Learning Semantic-Aware Representation in Visual-Language Models for Multi-Label Recognition with Partial Labels

Haoxian Ruan, Zhihua Xu, Zhijing Yang +3

Multi-label recognition with partial labels (MLR-PL), in which only some labels are known while others are unknown for each image, is a practical task in computer vision, since col…

cs.CV2020

Efficient Crowd Counting via Structured Knowledge Transfer

Lingbo Liu, Jiaqi Chen, Hefeng Wu +3

Crowd counting is an application-oriented task and its inference efficiency is crucial for real-world applications. However, most previous works relied on heavy backbone networks a…

cs.CV2025

Neural Scene Designer: Self-Styled Semantic Image Manipulation

Jianman Lin, Tianshui Chen, Chunmei Qing +4

Maintaining stylistic consistency is crucial for the cohesion and aesthetic appeal of images, a fundamental requirement in effective image editing and inpainting. However, existing…

cs.CV2026

Agentic Designer: Progressive Multi-Agent Collaboration for Structure-Aware Interior Layout Generation

Zhijing Yang, Haocheng Lin, Zhihua Xu +4

Generating realistic interior furniture layouts that strictly adhere to architectural constraints (e.g., walls, doors, and windows) remains a fundamental challenge in automated spa…

cs.CV2025

Exploiting Temporal Audio-Visual Correlation Embedding for Audio-Driven One-Shot Talking Head Animation

Zhihua Xu, Tianshui Chen, Zhijing Yang +3

The paramount challenge in audio-driven One-shot Talking Head Animation (ADOS-THA) lies in capturing subtle imperceptible changes between adjacent video frames. Inherently, the tem…

cs.CV2024

Heterogeneous Semantic Transfer for Multi-label Recognition with Partial Labels

Tianshui Chen, Tao Pu, Lingbo Liu +3

Multi-label image recognition with partial labels (MLR-PL), in which some labels are known while others are unknown for each image, may greatly reduce the cost of annotation and th…

cs.CV2018

Deep Reasoning with Knowledge Graph for Social Relationship Understanding

Zhouxia Wang, Tianshui Chen, Jimmy Ren +3

Social relationships (e.g., friends, couple etc.) form the basis of the social network in our daily life. Automatically interpreting such relationships bears a great potential for…

cs.CV2017

Content-Adaptive Sketch Portrait Generation by Decompositional Representation Learning

Dongyu Zhang, Liang Lin, Tianshui Chen +3

Sketch portrait generation benefits a wide range of applications such as digital entertainment and law enforcement. Although plenty of efforts have been dedicated to this task, sev…

cs.CV2018

Neural Task Planning with And-Or Graph Representations

Tianshui Chen, Riquan Chen, Lin Nie +3

This paper focuses on semantic task planning, i.e., predicting a sequence of actions toward accomplishing a specific task under a certain scene, which is a new problem in computer…

cs.CV2018

Learning to Segment Object Candidates via Recursive Neural Networks

Tianshui Chen, Liang Lin, Xian Wu +2

To avoid the exhaustive search over locations and scales, current state-of-the-art object detection systems usually involve a crucial component generating a batch of candidate obje…

cs.CV2025

SQLNet: Scale-Modulated Query and Localization Network for Few-Shot Class-Agnostic Counting

Hefeng Wu, Yandong Chen, Lingbo Liu +3

The class-agnostic counting (CAC) task has recently been proposed to solve the problem of counting all objects of an arbitrary class with several exemplars given in the input image…

cs.CV2026

Webly Supervised Multi-Label Recognition: Evaluation Benchmark and Dual-Branch Multi-Label Contrastive Learning

Zhihua Xu, Zhijing Yang, Yufeng Yang +1

Training deep learning models with freely available web images can reduce their dependence on costly manual annotations. Although webly supervised learning has been widely studied…

cs.RO2026

Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation

Zijian Song, Qichang Li, Sihan Qin +4

The scarcity of large-scale robotic data has motivated the repurposing of foundation models from other modalities for policy learning. In this work, we introduce PhysGen (Learning…

cs.CV2023

Semantic Representation and Dependency Learning for Multi-Label Image Recognition

Tao Pu, Mingzhan Sun, Hefeng Wu +3

Recently many multi-label image recognition (MLR) works have made significant progress by introducing pre-trained object detection models to generate lots of proposals or utilizing…

cs.CV2019

Learning Semantic-Specific Graph Representation for Multi-Label Image Recognition

Tianshui Chen, Muxin Xu, Xiaolu Hui +2

Recognizing multiple labels of images is a practical and challenging task, and significant progress has been made by searching semantic-aware regions and modeling label dependency.…

cs.CV2018

Fine-Grained Representation Learning and Recognition by Exploiting Hierarchical Semantic Embedding

Tianshui Chen, Wenxi Wu, Yuefang Gao +3

Object categories inherently form a hierarchy with different levels of concept abstraction, especially for fine-grained categories. For example, birds (Aves) can be categorized acc…

cs.CV2024

Adaptive Global-Local Representation Learning and Selection for Cross-Domain Facial Expression Recognition

Yuefang Gao, Yuhao Xie, Zeke Zexi Hu +2

Domain shift poses a significant challenge in Cross-Domain Facial Expression Recognition (CD-FER) due to the distribution variation across different domains. Current works mainly f…

cs.CV2023

RestoreFormer++: Towards Real-World Blind Face Restoration from Undegraded Key-Value Pairs

Zhouxia Wang, Jiawei Zhang, Tianshui Chen +2

Blind face restoration aims at recovering high-quality face images from those with unknown degradations. Current algorithms mainly introduce priors to complement high-quality detai…

cs.CV2017

Learning a Wavelet-like Auto-Encoder to Accelerate Deep Neural Networks

Tianshui Chen, Liang Lin, Wangmeng Zuo +2

Accelerating deep neural networks (DNNs) has been attracting increasing attention as it can benefit a wide range of applications, e.g., enabling mobile systems with limited computi…

cs.CV2016

Character Proposal Network for Robust Text Extraction

Shuye Zhang, Mude Lin, Tianshui Chen +2

Maximally stable extremal regions (MSER), which is a popular method to generate character proposals/candidates, has shown superior performance in scene text detection. However, the…

cs.CV2023

Contrastive Transformer Learning with Proximity Data Generation for Text-Based Person Search

Hefeng Wu, Weifeng Chen, Zhibin Liu +3

Given a descriptive text query, text-based person search (TBPS) aims to retrieve the best-matched target person from an image gallery. Such a cross-modal retrieval task is quite ch…

cs.CV2024

Dynamic Correlation Learning and Regularization for Multi-Label Confidence Calibration

Tianshui Chen, Weihang Wang, Tao Pu +4

Modern visual recognition models often display overconfidence due to their reliance on complex deep neural networks and one-hot target supervision, resulting in unreliable confiden…

cs.CV2026

Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding

Yuan Xie, Tianshui Chen, Zheng Ge +1

Long-form video understanding, characterized by long-range temporal dependencies and multiple events, remains a challenge. Existing methods often rely on static reasoning or extern…

cs.CV2026

Neural Clothing Tryer: Customized Virtual Try-On via Semantic Enhancement and Controlling Diffusion Model

Zhijing Yang, Weiwei Zhang, Mingliang Yang +5

This work aims to address a novel Customized Virtual Try-ON (Cu-VTON) task, enabling the superimposition of a specified garment onto a model that can be customized in terms of appe…

cs.CV2026

Geometry-Editable and Appearance-Preserving Object Compositon

Jianman Lin, Haojie Li, Chunmei Qing +3

General object composition (GOC) aims to seamlessly integrate a target object into a background scene with desired geometric properties, while simultaneously preserving its fine-gr…

cs.CV2025

Physical Autoregressive Model for Robotic Manipulation without Action Pretraining

Zijian Song, Sihan Qin, Tianshui Chen +2

The scarcity of manipulation data has motivated the use of pretrained large models from other modalities in robotics. In this work, we build upon autoregressive video generation mo…

cs.CV2024

Diff-Mosaic: Augmenting Realistic Representations in Infrared Small Target Detection via Diffusion Prior

Yukai Shi, Yupei Lin, Pengxu Wei +3

Recently, researchers have proposed various deep learning methods to accurately detect infrared targets with the characteristics of indistinct shape and texture. Due to the limited…

cs.CV2023

Spatial-Temporal Knowledge-Embedded Transformer for Video Scene Graph Generation

Tao Pu, Tianshui Chen, Hefeng Wu +2

Video scene graph generation (VidSGG) aims to identify objects in visual scenes and infer their relationships for a given video. It requires not only a comprehensive understanding…

cs.CV2019

Semi-Supervised Video Salient Object Detection Using Pseudo-Labels

Pengxiang Yan, Guanbin Li, Yuan Xie +4

Deep learning-based video salient object detection has recently achieved great success with its performance significantly outperforming any other unsupervised methods. However, exi…

cs.CV2020

Knowledge-Guided Multi-Label Few-Shot Learning for General Image Recognition

Tianshui Chen, Liang Lin, Riquan Chen +2

Recognizing multiple labels of an image is a practical yet challenging task, and remarkable progress has been achieved by searching for semantic regions and exploiting label depend…

cs.CV2024

Dual-Perspective Semantic-Aware Representation Blending for Multi-Label Image Recognition with Partial Labels

Tao Pu, Tianshui Chen, Hefeng Wu +3

Despite achieving impressive progress, current multi-label image recognition (MLR) algorithms heavily depend on large-scale datasets with complete labels, making collecting large-s…

cs.CV2023

Mimic: Speaking Style Disentanglement for Speech-Driven 3D Facial Animation

Hui Fu, Zeqing Wang, Ke Gong +5

Speech-driven 3D facial animation aims to synthesize vivid facial animations that accurately synchronize with speech and match the unique speaking style. However, existing works pr…

cs.CV2024

Analysis and Benchmarking of Extending Blind Face Image Restoration to Videos

Zhouxia Wang, Jiawei Zhang, Xintao Wang +4

Recent progress in blind face restoration has resulted in producing high-quality restored results for static images. However, efforts to extend these advancements to video scenario…

cs.RO2026

RADAR: Benchmarking Vision-Language-Action Generalization via Real-World Dynamics, Spatial-Physical Intelligence, and Autonomous Evaluation

Yuhao Chen, Zhihao Zhan, Xiaoxin Lin +11

VLA models have achieved remarkable progress in embodied intelligence; however, their evaluation remains largely confined to simulations or highly constrained real-world settings.…

cs.CV2025

Monocular and Generalizable Gaussian Talking Head Animation

Shengjie Gong, Haojie Li, Jiapeng Tang +5

In this work, we introduce Monocular and Generalizable Gaussian Talking Head Animation (MGGTalk), which requires monocular datasets and generalizes to unseen identities without per…

cs.CV2022

Structured Semantic Transfer for Multi-Label Recognition with Partial Labels

Tianshui Chen, Tao Pu, Hefeng Wu +2

Multi-label image recognition is a fundamental yet practical task because real-world images inherently possess multiple semantic labels. However, it is difficult to collect large-s…

cs.CV2022

Aerial Images Meet Crowdsourced Trajectories: A New Approach to Robust Road Extraction

Lingbo Liu, Zewei Yang, Guanbin Li +3

Land remote sensing analysis is a crucial research in earth science. In this work, we focus on a challenging task of land analysis, i.e., automatic extraction of traffic roads from…

cs.CV2023

Perception and Semantic Aware Regularization for Sequential Confidence Calibration

Zhenghua Peng, Yu Luo, Tianshui Chen +2

Deep sequence recognition (DSR) models receive increasing attention due to their superior application to various applications. Most DSR models use merely the target sequences as su…

cs.CV2022

Exploring Negatives in Contrastive Learning for Unpaired Image-to-Image Translation

Yupei Lin, Sen Zhang, Tianshui Chen +3

Unpaired image-to-image translation aims to find a mapping between the source domain and the target domain. To alleviate the problem of the lack of supervised labels for the source…

cs.CV2019

Knowledge-Embedded Routing Network for Scene Graph Generation

Tianshui Chen, Weihao Yu, Riquan Chen +1

To understand a scene in depth not only involves locating/recognizing individual objects, but also requires to infer the relationships and interactions among them. However, since t…

cs.CV2017

Knowledge-Guided Recurrent Neural Network Learning for Task-Oriented Action Prediction

Liang Lin, Lili Huang, Tianshui Chen +2

This paper aims at task-oriented action prediction, i.e., predicting a sequence of actions towards accomplishing a specific task under a certain scene, which is a new problem in co…

cs.CV2020

Fine-Grained Image Captioning with Global-Local Discriminative Objective

Jie Wu, Tianshui Chen, Hefeng Wu +3

Significant progress has been made in recent years in image captioning, an active topic in the fields of vision and language. However, existing methods tend to yield overly general…

cs.CV2023

Open-World Pose Transfer via Sequential Test-Time Adaption

Junyang Chen, Xiaoyu Xian, Zhijing Yang +5

Pose transfer aims to transfer a given person into a specified posture, has recently attracted considerable attention. A typical pose transfer framework usually employs representat…

cs.RO2026

E0: Enhancing Generalization and Fine-Grained Control in VLA Models via Tweedie Discrete Diffusion

Zhihao Zhan, Jiaying Zhou, Likui Zhang +10

Vision-Language-Action (VLA) models offer a unified framework for robotic manipulation by integrating visual perception, language understanding, and control generation. However, ex…

cs.CV2025

Learning Semantic-Aware Threshold for Multi-Label Image Recognition with Partial Labels

Haoxian Ruan, Zhihua Xu, Zhijing Yang +4

Multi-label image recognition with partial labels (MLR-PL) is designed to train models using a mix of known and unknown labels. Traditional methods rely on semantic or feature corr…

cs.CV2025

Contrastive Decoupled Representation Learning and Regularization for Speech-Preserving Facial Expression Manipulation

Tianshui Chen, Jianman Lin, Zhijing Yang +3

Speech-preserving facial expression manipulation (SPFEM) aims to modify a talking head to display a specific reference emotion while preserving the mouth animation of source spoken…

cs.AI2025

Towards CausalGPT: A Multi-Agent Approach for Faithful Knowledge Reasoning via Promoting Causal Consistency in LLMs

Ziyi Tang, Ruilin Wang, Weixing Chen +6

Despite the progress of foundation models, knowledge-based reasoning remains a persistent challenge due to their limited capacity for knowledge recall and inference. Existing metho…

cs.AI2026

Learning Hierarchical and Geometry-Aware Graph Representations for Text-to-CAD

Shengjie Gong, Wenjie Peng, Hongyuan Chen +5

Text-to-CAD code generation is a long-horizon task that translates textual instructions into long sequences of interdependent operations. Existing methods typically decode text dir…

cs.CV2021

AU-Expression Knowledge Constrained Representation Learning for Facial Expression Recognition

Tao Pu, Tianshui Chen, Yuan Xie +2

Recognizing human emotion/expressions automatically is quite an expected ability for intelligent robotics, as it can promote better communication and cooperation with humans. Curre…

cs.CV2021

Cross-Domain Facial Expression Recognition: A Unified Evaluation Benchmark and Adversarial Graph Learning

Tianshui Chen, Tao Pu, Hefeng Wu +3

To address the problem of data inconsistencies among different facial expression recognition (FER) datasets, many cross-domain FER methods (CD-FERs) have been extensively devised i…

cs.CV2024

MotionCtrl: A Unified and Flexible Motion Controller for Video Generation

Zhouxia Wang, Ziyang Yuan, Xintao Wang +4

Motions in a video primarily consist of camera motion, induced by camera movement, and object motion, resulting from object movement. Accurate control of both camera and object mot…

cs.CV2020

Adversarial Graph Representation Adaptation for Cross-Domain Facial Expression Recognition

Yuan Xie, Tianshui Chen, Tao Pu +2

Data inconsistency and bias are inevitable among different facial expression recognition (FER) datasets due to subjective annotating process and different collecting conditions. Re…

cs.CV2015

DISC: Deep Image Saliency Computing via Progressive Representation Learning

Tianshui Chen, Liang Lin, Lingbo Liu +2

Salient object detection increasingly receives attention as an important component or step in several pattern recognition and image processing tasks. Although a variety of powerful…

cs.CV2025

ReplayCAD: Generative Diffusion Replay for Continual Anomaly Detection

Lei Hu, Zhiyong Gan, Ling Deng +4

Continual Anomaly Detection (CAD) enables anomaly detection models in learning new classes while preserving knowledge of historical classes. CAD faces two key challenges: catastrop…

cs.CV2017

Recurrent Attentional Reinforcement Learning for Multi-label Image Recognition

Tianshui Chen, Zhouxia Wang, Guanbin Li +1

Recognizing multiple labels of images is a fundamental but challenging task in computer vision, and remarkable progress has been attained by localizing semantic-aware image regions…

cs.CV2026

Learning Spatial-Temporal Coherent Correlations for Speech-Preserving Facial Expression Manipulation

Tianshui Chen, Jianman Lin, Zhijing Yang +3

Speech-preserving facial expression manipulation (SPFEM) aims to modify facial emotions while meticulously maintaining the mouth animation associated with spoken content. Current w…

cs.CV2025

In-Situ Tweedie Discrete Diffusion Models

Xiao Li, Jiaqi Zhang, Shuxiang Zhang +3

While diffusion models excel at generating continuous data such as images, adapting them to discrete tasks has relied on indirect approaches that either operate in continuous embed…

cs.CV2022

Category-Adaptive Label Discovery and Noise Rejection for Multi-label Image Recognition with Partial Positive Labels

Tao Pu, Qianru Lao, Hefeng Wu +2

As a promising solution of reducing annotation cost, training multi-label models with partial positive labels (MLR-PPL), in which merely few positive labels are known while other a…

cs.CV2026

Exploring Talking Head Models With Adjacent Frame Prior for Speech-Preserving Facial Expression Manipulation

Zhenxuan Lu, Zhihua Xu, Zhijing Yang +4

Speech-Preserving Facial Expression Manipulation (SPFEM) is an innovative technique aimed at altering facial expressions in images and videos while retaining the original mouth mov…