Publications (67)
Robotic Manipulation is Vision-to-Geometry Mapping (): Vision-Geometry Backbones over Language and Video Models
Zijian Song, Qichang Li, Jiawei Zhou +4
At its core, robotic manipulation is a problem of vision-to-geometry mapping (). Physical actions are fundamentally defined by geometric properties like 3D posi…
Knowledge-Embedded Representation Learning for Fine-Grained Image Recognition
Tianshui Chen, Liang Lin, Riquan Chen +2
Humans can naturally understand an image in depth with the aid of rich knowledge accumulated from daily lives or professions. For example, to achieve fine-grained image recognition…
Globally Correlation-Aware Hard Negative Generation
Wenjie Peng, Hongxiang Huang, Tianshui Chen +3
Hard negative generation aims to generate informative negative samples that help to determine the decision boundaries and thus facilitate advancing deep metric learning. Current wo…
Personalized Cross-Modal Emotional Correlation Learning for Speech-Preserving Facial Expression Manipulation
Tianshui Chen, Yujie Zhu, Jianman Lin +4
Speech-preserving facial expression manipulation (SPFEM) aims to enhance human expressiveness without altering mouth movements tied to the original speech. A primary challenge in t…
Knowledge Graph Transfer Network for Few-Shot Recognition
Riquan Chen, Tianshui Chen, Xiaolu Hui +3
Few-shot learning aims to learn novel categories from very few samples given some base categories with sufficient training samples. The main challenge of this task is the novel cat…
OccluMix: Towards De-Occlusion Virtual Try-on by Semantically-Guided Mixup
Zhijing Yang, Junyang Chen, Yukai Shi +3
Image Virtual try-on aims at replacing the cloth on a personal image with a garment image (in-shop clothes), which has attracted increasing attention from the multimedia and comput…
Multi-label Image Recognition by Recurrently Discovering Attentional Regions
Zhouxia Wang, Tianshui Chen, Guanbin Li +2
This paper proposes a novel deep architecture to address multi-label image recognition, a fundamental and practical task towards general visual understanding. Current solutions for…
Semantic-Aware Representation Blending for Multi-Label Image Recognition with Partial Labels
Tao Pu, Tianshui Chen, Hefeng Wu +1
Training the multi-label image recognition models with partial labels, in which merely some labels are known while others are unknown for each image, is a considerably challenging…
Learning Semantic-Aware Representation in Visual-Language Models for Multi-Label Recognition with Partial Labels
Haoxian Ruan, Zhihua Xu, Zhijing Yang +3
Multi-label recognition with partial labels (MLR-PL), in which only some labels are known while others are unknown for each image, is a practical task in computer vision, since col…
Efficient Crowd Counting via Structured Knowledge Transfer
Lingbo Liu, Jiaqi Chen, Hefeng Wu +3
Crowd counting is an application-oriented task and its inference efficiency is crucial for real-world applications. However, most previous works relied on heavy backbone networks a…
Neural Scene Designer: Self-Styled Semantic Image Manipulation
Jianman Lin, Tianshui Chen, Chunmei Qing +4
Maintaining stylistic consistency is crucial for the cohesion and aesthetic appeal of images, a fundamental requirement in effective image editing and inpainting. However, existing…
Agentic Designer: Progressive Multi-Agent Collaboration for Structure-Aware Interior Layout Generation
Zhijing Yang, Haocheng Lin, Zhihua Xu +4
Generating realistic interior furniture layouts that strictly adhere to architectural constraints (e.g., walls, doors, and windows) remains a fundamental challenge in automated spa…
Exploiting Temporal Audio-Visual Correlation Embedding for Audio-Driven One-Shot Talking Head Animation
Zhihua Xu, Tianshui Chen, Zhijing Yang +3
The paramount challenge in audio-driven One-shot Talking Head Animation (ADOS-THA) lies in capturing subtle imperceptible changes between adjacent video frames. Inherently, the tem…
Heterogeneous Semantic Transfer for Multi-label Recognition with Partial Labels
Tianshui Chen, Tao Pu, Lingbo Liu +3
Multi-label image recognition with partial labels (MLR-PL), in which some labels are known while others are unknown for each image, may greatly reduce the cost of annotation and th…
Deep Reasoning with Knowledge Graph for Social Relationship Understanding
Zhouxia Wang, Tianshui Chen, Jimmy Ren +3
Social relationships (e.g., friends, couple etc.) form the basis of the social network in our daily life. Automatically interpreting such relationships bears a great potential for…
Content-Adaptive Sketch Portrait Generation by Decompositional Representation Learning
Dongyu Zhang, Liang Lin, Tianshui Chen +3
Sketch portrait generation benefits a wide range of applications such as digital entertainment and law enforcement. Although plenty of efforts have been dedicated to this task, sev…
Neural Task Planning with And-Or Graph Representations
Tianshui Chen, Riquan Chen, Lin Nie +3
This paper focuses on semantic task planning, i.e., predicting a sequence of actions toward accomplishing a specific task under a certain scene, which is a new problem in computer…
Learning to Segment Object Candidates via Recursive Neural Networks
Tianshui Chen, Liang Lin, Xian Wu +2
To avoid the exhaustive search over locations and scales, current state-of-the-art object detection systems usually involve a crucial component generating a batch of candidate obje…
SQLNet: Scale-Modulated Query and Localization Network for Few-Shot Class-Agnostic Counting
Hefeng Wu, Yandong Chen, Lingbo Liu +3
The class-agnostic counting (CAC) task has recently been proposed to solve the problem of counting all objects of an arbitrary class with several exemplars given in the input image…
Webly Supervised Multi-Label Recognition: Evaluation Benchmark and Dual-Branch Multi-Label Contrastive Learning
Zhihua Xu, Zhijing Yang, Yufeng Yang +1
Training deep learning models with freely available web images can reduce their dependence on costly manual annotations. Although webly supervised learning has been widely studied…
Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation
Zijian Song, Qichang Li, Sihan Qin +4
The scarcity of large-scale robotic data has motivated the repurposing of foundation models from other modalities for policy learning. In this work, we introduce PhysGen (Learning…
Semantic Representation and Dependency Learning for Multi-Label Image Recognition
Tao Pu, Mingzhan Sun, Hefeng Wu +3
Recently many multi-label image recognition (MLR) works have made significant progress by introducing pre-trained object detection models to generate lots of proposals or utilizing…
Learning Semantic-Specific Graph Representation for Multi-Label Image Recognition
Tianshui Chen, Muxin Xu, Xiaolu Hui +2
Recognizing multiple labels of images is a practical and challenging task, and significant progress has been made by searching semantic-aware regions and modeling label dependency.…
Fine-Grained Representation Learning and Recognition by Exploiting Hierarchical Semantic Embedding
Tianshui Chen, Wenxi Wu, Yuefang Gao +3
Object categories inherently form a hierarchy with different levels of concept abstraction, especially for fine-grained categories. For example, birds (Aves) can be categorized acc…
Adaptive Global-Local Representation Learning and Selection for Cross-Domain Facial Expression Recognition
Yuefang Gao, Yuhao Xie, Zeke Zexi Hu +2
Domain shift poses a significant challenge in Cross-Domain Facial Expression Recognition (CD-FER) due to the distribution variation across different domains. Current works mainly f…
RestoreFormer++: Towards Real-World Blind Face Restoration from Undegraded Key-Value Pairs
Zhouxia Wang, Jiawei Zhang, Tianshui Chen +2
Blind face restoration aims at recovering high-quality face images from those with unknown degradations. Current algorithms mainly introduce priors to complement high-quality detai…
Learning a Wavelet-like Auto-Encoder to Accelerate Deep Neural Networks
Tianshui Chen, Liang Lin, Wangmeng Zuo +2
Accelerating deep neural networks (DNNs) has been attracting increasing attention as it can benefit a wide range of applications, e.g., enabling mobile systems with limited computi…
Character Proposal Network for Robust Text Extraction
Shuye Zhang, Mude Lin, Tianshui Chen +2
Maximally stable extremal regions (MSER), which is a popular method to generate character proposals/candidates, has shown superior performance in scene text detection. However, the…
Contrastive Transformer Learning with Proximity Data Generation for Text-Based Person Search
Hefeng Wu, Weifeng Chen, Zhibin Liu +3
Given a descriptive text query, text-based person search (TBPS) aims to retrieve the best-matched target person from an image gallery. Such a cross-modal retrieval task is quite ch…
Dynamic Correlation Learning and Regularization for Multi-Label Confidence Calibration
Tianshui Chen, Weihang Wang, Tao Pu +4
Modern visual recognition models often display overconfidence due to their reliance on complex deep neural networks and one-hot target supervision, resulting in unreliable confiden…
Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding
Yuan Xie, Tianshui Chen, Zheng Ge +1
Long-form video understanding, characterized by long-range temporal dependencies and multiple events, remains a challenge. Existing methods often rely on static reasoning or extern…
Neural Clothing Tryer: Customized Virtual Try-On via Semantic Enhancement and Controlling Diffusion Model
Zhijing Yang, Weiwei Zhang, Mingliang Yang +5
This work aims to address a novel Customized Virtual Try-ON (Cu-VTON) task, enabling the superimposition of a specified garment onto a model that can be customized in terms of appe…
Geometry-Editable and Appearance-Preserving Object Compositon
Jianman Lin, Haojie Li, Chunmei Qing +3
General object composition (GOC) aims to seamlessly integrate a target object into a background scene with desired geometric properties, while simultaneously preserving its fine-gr…
Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
Zijian Song, Sihan Qin, Tianshui Chen +2
The scarcity of manipulation data has motivated the use of pretrained large models from other modalities in robotics. In this work, we build upon autoregressive video generation mo…
Diff-Mosaic: Augmenting Realistic Representations in Infrared Small Target Detection via Diffusion Prior
Yukai Shi, Yupei Lin, Pengxu Wei +3
Recently, researchers have proposed various deep learning methods to accurately detect infrared targets with the characteristics of indistinct shape and texture. Due to the limited…
Spatial-Temporal Knowledge-Embedded Transformer for Video Scene Graph Generation
Tao Pu, Tianshui Chen, Hefeng Wu +2
Video scene graph generation (VidSGG) aims to identify objects in visual scenes and infer their relationships for a given video. It requires not only a comprehensive understanding…
Semi-Supervised Video Salient Object Detection Using Pseudo-Labels
Pengxiang Yan, Guanbin Li, Yuan Xie +4
Deep learning-based video salient object detection has recently achieved great success with its performance significantly outperforming any other unsupervised methods. However, exi…
Knowledge-Guided Multi-Label Few-Shot Learning for General Image Recognition
Tianshui Chen, Liang Lin, Riquan Chen +2
Recognizing multiple labels of an image is a practical yet challenging task, and remarkable progress has been achieved by searching for semantic regions and exploiting label depend…
Dual-Perspective Semantic-Aware Representation Blending for Multi-Label Image Recognition with Partial Labels
Tao Pu, Tianshui Chen, Hefeng Wu +3
Despite achieving impressive progress, current multi-label image recognition (MLR) algorithms heavily depend on large-scale datasets with complete labels, making collecting large-s…
Mimic: Speaking Style Disentanglement for Speech-Driven 3D Facial Animation
Hui Fu, Zeqing Wang, Ke Gong +5
Speech-driven 3D facial animation aims to synthesize vivid facial animations that accurately synchronize with speech and match the unique speaking style. However, existing works pr…
Analysis and Benchmarking of Extending Blind Face Image Restoration to Videos
Zhouxia Wang, Jiawei Zhang, Xintao Wang +4
Recent progress in blind face restoration has resulted in producing high-quality restored results for static images. However, efforts to extend these advancements to video scenario…
RADAR: Benchmarking Vision-Language-Action Generalization via Real-World Dynamics, Spatial-Physical Intelligence, and Autonomous Evaluation
Yuhao Chen, Zhihao Zhan, Xiaoxin Lin +11
VLA models have achieved remarkable progress in embodied intelligence; however, their evaluation remains largely confined to simulations or highly constrained real-world settings.…
Monocular and Generalizable Gaussian Talking Head Animation
Shengjie Gong, Haojie Li, Jiapeng Tang +5
In this work, we introduce Monocular and Generalizable Gaussian Talking Head Animation (MGGTalk), which requires monocular datasets and generalizes to unseen identities without per…
Structured Semantic Transfer for Multi-Label Recognition with Partial Labels
Tianshui Chen, Tao Pu, Hefeng Wu +2
Multi-label image recognition is a fundamental yet practical task because real-world images inherently possess multiple semantic labels. However, it is difficult to collect large-s…
Aerial Images Meet Crowdsourced Trajectories: A New Approach to Robust Road Extraction
Lingbo Liu, Zewei Yang, Guanbin Li +3
Land remote sensing analysis is a crucial research in earth science. In this work, we focus on a challenging task of land analysis, i.e., automatic extraction of traffic roads from…
Perception and Semantic Aware Regularization for Sequential Confidence Calibration
Zhenghua Peng, Yu Luo, Tianshui Chen +2
Deep sequence recognition (DSR) models receive increasing attention due to their superior application to various applications. Most DSR models use merely the target sequences as su…
Exploring Negatives in Contrastive Learning for Unpaired Image-to-Image Translation
Yupei Lin, Sen Zhang, Tianshui Chen +3
Unpaired image-to-image translation aims to find a mapping between the source domain and the target domain. To alleviate the problem of the lack of supervised labels for the source…
Knowledge-Embedded Routing Network for Scene Graph Generation
Tianshui Chen, Weihao Yu, Riquan Chen +1
To understand a scene in depth not only involves locating/recognizing individual objects, but also requires to infer the relationships and interactions among them. However, since t…
Knowledge-Guided Recurrent Neural Network Learning for Task-Oriented Action Prediction
Liang Lin, Lili Huang, Tianshui Chen +2
This paper aims at task-oriented action prediction, i.e., predicting a sequence of actions towards accomplishing a specific task under a certain scene, which is a new problem in co…
Fine-Grained Image Captioning with Global-Local Discriminative Objective
Jie Wu, Tianshui Chen, Hefeng Wu +3
Significant progress has been made in recent years in image captioning, an active topic in the fields of vision and language. However, existing methods tend to yield overly general…
Open-World Pose Transfer via Sequential Test-Time Adaption
Junyang Chen, Xiaoyu Xian, Zhijing Yang +5
Pose transfer aims to transfer a given person into a specified posture, has recently attracted considerable attention. A typical pose transfer framework usually employs representat…
E0: Enhancing Generalization and Fine-Grained Control in VLA Models via Tweedie Discrete Diffusion
Zhihao Zhan, Jiaying Zhou, Likui Zhang +10
Vision-Language-Action (VLA) models offer a unified framework for robotic manipulation by integrating visual perception, language understanding, and control generation. However, ex…
Learning Semantic-Aware Threshold for Multi-Label Image Recognition with Partial Labels
Haoxian Ruan, Zhihua Xu, Zhijing Yang +4
Multi-label image recognition with partial labels (MLR-PL) is designed to train models using a mix of known and unknown labels. Traditional methods rely on semantic or feature corr…
Contrastive Decoupled Representation Learning and Regularization for Speech-Preserving Facial Expression Manipulation
Tianshui Chen, Jianman Lin, Zhijing Yang +3
Speech-preserving facial expression manipulation (SPFEM) aims to modify a talking head to display a specific reference emotion while preserving the mouth animation of source spoken…
Towards CausalGPT: A Multi-Agent Approach for Faithful Knowledge Reasoning via Promoting Causal Consistency in LLMs
Ziyi Tang, Ruilin Wang, Weixing Chen +6
Despite the progress of foundation models, knowledge-based reasoning remains a persistent challenge due to their limited capacity for knowledge recall and inference. Existing metho…
Learning Hierarchical and Geometry-Aware Graph Representations for Text-to-CAD
Shengjie Gong, Wenjie Peng, Hongyuan Chen +5
Text-to-CAD code generation is a long-horizon task that translates textual instructions into long sequences of interdependent operations. Existing methods typically decode text dir…
AU-Expression Knowledge Constrained Representation Learning for Facial Expression Recognition
Tao Pu, Tianshui Chen, Yuan Xie +2
Recognizing human emotion/expressions automatically is quite an expected ability for intelligent robotics, as it can promote better communication and cooperation with humans. Curre…
Cross-Domain Facial Expression Recognition: A Unified Evaluation Benchmark and Adversarial Graph Learning
Tianshui Chen, Tao Pu, Hefeng Wu +3
To address the problem of data inconsistencies among different facial expression recognition (FER) datasets, many cross-domain FER methods (CD-FERs) have been extensively devised i…
MotionCtrl: A Unified and Flexible Motion Controller for Video Generation
Zhouxia Wang, Ziyang Yuan, Xintao Wang +4
Motions in a video primarily consist of camera motion, induced by camera movement, and object motion, resulting from object movement. Accurate control of both camera and object mot…
Adversarial Graph Representation Adaptation for Cross-Domain Facial Expression Recognition
Yuan Xie, Tianshui Chen, Tao Pu +2
Data inconsistency and bias are inevitable among different facial expression recognition (FER) datasets due to subjective annotating process and different collecting conditions. Re…
DISC: Deep Image Saliency Computing via Progressive Representation Learning
Tianshui Chen, Liang Lin, Lingbo Liu +2
Salient object detection increasingly receives attention as an important component or step in several pattern recognition and image processing tasks. Although a variety of powerful…
ReplayCAD: Generative Diffusion Replay for Continual Anomaly Detection
Lei Hu, Zhiyong Gan, Ling Deng +4
Continual Anomaly Detection (CAD) enables anomaly detection models in learning new classes while preserving knowledge of historical classes. CAD faces two key challenges: catastrop…
Recurrent Attentional Reinforcement Learning for Multi-label Image Recognition
Tianshui Chen, Zhouxia Wang, Guanbin Li +1
Recognizing multiple labels of images is a fundamental but challenging task in computer vision, and remarkable progress has been attained by localizing semantic-aware image regions…
Learning Spatial-Temporal Coherent Correlations for Speech-Preserving Facial Expression Manipulation
Tianshui Chen, Jianman Lin, Zhijing Yang +3
Speech-preserving facial expression manipulation (SPFEM) aims to modify facial emotions while meticulously maintaining the mouth animation associated with spoken content. Current w…
In-Situ Tweedie Discrete Diffusion Models
Xiao Li, Jiaqi Zhang, Shuxiang Zhang +3
While diffusion models excel at generating continuous data such as images, adapting them to discrete tasks has relied on indirect approaches that either operate in continuous embed…
Category-Adaptive Label Discovery and Noise Rejection for Multi-label Image Recognition with Partial Positive Labels
Tao Pu, Qianru Lao, Hefeng Wu +2
As a promising solution of reducing annotation cost, training multi-label models with partial positive labels (MLR-PPL), in which merely few positive labels are known while other a…
Exploring Talking Head Models With Adjacent Frame Prior for Speech-Preserving Facial Expression Manipulation
Zhenxuan Lu, Zhihua Xu, Zhijing Yang +4
Speech-Preserving Facial Expression Manipulation (SPFEM) is an innovative technique aimed at altering facial expressions in images and videos while retaining the original mouth mov…