Publications (262)
Neural Pose Transfer by Spatially Adaptive Instance Normalization
Jiashun Wang, Chao Wen, Yanwei Fu +4
Pose transfer has been studied for decades, in which the pose of a source mesh is applied to a target mesh. Particularly in this paper, we are interested in transferring the pose o…
Polaris: Open-ended Interactive Robotic Manipulation via Syn2Real Visual Grounding and Large Language Models
Tianyu Wang, Haitao Lin, Junqiu Yu +1
This paper investigates the task of the open-ended interactive robotic manipulation on table-top scenarios. While recent Large Language Models (LLMs) enhance robots' comprehension…
InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
Xinyi Chen, Yilun Chen, Yanwei Fu +26
We introduce InternVLA-M1, a unified framework for spatial grounding and robot control that advances instruction-following robots toward scalable, general-purpose intelligence. Its…
Towards Reliable and Holistic Visual In-Context Learning Prompt Selection
Wenxiao Wu, Jing-Hao Xue, Chengming Xu +5
Visual In-Context Learning (VICL) has emerged as a prominent approach for adapting visual foundation models to novel tasks, by effectively exploiting contextual information embedde…
A Jointly Learned Deep Architecture for Facial Attribute Analysis and Face Detection in the Wild
Keke He, Yanwei Fu, Xiangyang Xue
Facial attribute analysis in the real world scenario is very challenging mainly because of complex face variations. Existing works of analyzing face attributes are mostly based on…
Recent Advances in Zero-shot Recognition
Yanwei Fu, Tao Xiang, Yu-Gang Jiang +3
With the recent renaissance of deep convolution neural networks, encouraging breakthroughs have been achieved on the supervised recognition tasks, where each class has sufficient t…
Semi-Latent GAN: Learning to generate and modify facial images from attributes
Weidong Yin, Yanwei Fu, Leonid Sigal +1
Generating and manipulating human facial images using high-level attributal controls are important and interesting problems. The models proposed in previous work can solve one of t…
ImpDet: Exploring Implicit Fields for 3D Object Detection
Xuelin Qian, Li Wang, Yi Zhu +3
Conventional 3D object detection approaches concentrate on bounding boxes representation learning with several parameters, i.e., localization, dimension, and orientation. Despite i…
Image Deformation Meta-Networks for One-Shot Learning
Zitian Chen, Yanwei Fu, Yu-Xiong Wang +3
Humans can robustly learn novel visual concepts even when images undergo various deformations and lose certain information. Mimicking the same behavior and synthesizing deformed in…
ST4VLA: Spatially Guided Training for Vision-Language-Action Models
Jinhui Ye, Fangjing Wang, Ning Gao +9
Large vision-language models (VLMs) excel at multimodal understanding but fall short when extended to embodied tasks, where instructions must be transformed into low-level motor ac…
Transductive Multi-label Zero-shot Learning
Yanwei Fu, Yongxin Yang, Tim Hospedales +2
Zero-shot learning has received increasing interest as a means to alleviate the often prohibitive expense of annotating training data for large scale recognition problems. These me…
FitDiT: Advancing the Authentic Garment Details for High-fidelity Virtual Try-on
Boyuan Jiang, Xiaobin Hu, Donghao Luo +7
Although image-based virtual try-on has made considerable progress, emerging approaches still encounter challenges in producing high-fidelity and robust fitting images across diver…
EgoSound: Benchmarking Sound Understanding in Egocentric Videos
Bingwen Zhu, Yuqian Fu, Qiaole Dong +6
Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in vision-language understanding. Yet, human perception is inherently multisensory, integrating…
ArtWeaver: Advanced Dynamic Style Integration via Diffusion Model
Chengming Xu, Kai Hu, Qilin Wang +5
Stylized Text-to-Image Generation (STIG) aims to generate images from text prompts and style reference images. In this paper, we present ArtWeaver, a novel framework that leverages…
Delving into Data: Effectively Substitute Training for Black-box Attack
Wenxuan Wang, Bangjie Yin, Taiping Yao +6
Deep models have shown their vulnerability when processing adversarial samples. As for the black-box attack, without access to the architecture and weights of the attacked model, t…
Stacked Semantic-Guided Attention Model for Fine-Grained Zero-Shot Learning
Yunlong Yu, Zhong Ji, Yanwei Fu +3
Zero-Shot Learning (ZSL) is achieved via aligning the semantic relationships between the global image feature vector and the corresponding class semantic descriptions. However, usi…
CustAny: Customizing Anything from A Single Example
Lingjie Kong, Kai Wu, Xiaobin Hu +8
Recent advances in diffusion-based text-to-image models have simplified creating high-fidelity images, but preserving the identity (ID) of specific elements, like a personal dog, i…
NSFW-Classifier Guided Prompt Sanitization for Safe Text-to-Image Generation
Yu Xie, Chengjie Zeng, Lingyun Zhang +1
The rapid advancement of text-to-image (T2I) models, such as Stable Diffusion, has enhanced their capability to synthesize images from textual prompts. However, this progress also…
Sketch-BERT: Learning Sketch Bidirectional Encoder Representation from Transformers by Self-supervised Learning of Sketch Gestalt
Hangyu Lin, Yanwei Fu, Yu-Gang Jiang +1
Previous researches of sketches often considered sketches in pixel format and leveraged CNN based models in the sketch understanding. Fundamentally, a sketch is stored as a sequenc…
StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering
Ming Xie, Zizheng Huang, Xudong Tan +6
While streaming omni-video understanding demands continuous perception and proactive, real-time interaction, this crucial area remains largely under-explored. Current omni-modal me…
PPMStereo: Pick-and-Play Memory Construction for Consistent Dynamic Stereo Matching
Yun Wang, Junjie Hu, Qiaole Dong +4
Temporally consistent depth estimation from stereo video is critical for real-world applications such as augmented reality, where inconsistent depth estimation disrupts the immersi…
Aligned Stable Inpainting: Mitigating Unwanted Object Insertion and Preserving Color Consistency
Yikai Wang, Junqiu Yu, Chenjie Cao +2
Generative image inpainting can produce realistic results even with large, irregular masks, but existing methods still suffer from two common problems: (1) Unwanted object insertio…
Local Slot Attention for Vision-and-Language Navigation
Yifeng Zhuang, Qiang Sun, Yanwei Fu +2
Vision-and-language navigation (VLN), a frontier study aiming to pave the way for general-purpose robots, has been a hot topic in the computer vision and natural language processin…
TP-MDDN: Task-Preferenced Multi-Demand-Driven Navigation with Autonomous Decision-Making
Shanshan Li, Da Huang, Yu He +3
In daily life, people often move through spaces to find objects that meet their needs, posing a key challenge in embodied AI. Traditional Demand-Driven Navigation (DDN) handles one…
SparseGrasp: Robotic Grasping via 3D Semantic Gaussian Splatting from Sparse Multi-View RGB Images
Junqiu Yu, Xinlin Ren, Yongchong Gu +7
Language-guided robotic grasping is a rapidly advancing field where robots are instructed using human language to grasp specific objects. However, existing methods often depend on…
Deep Learning for Video Classification and Captioning
Zuxuan Wu, Ting Yao, Yanwei Fu +1
Accelerated by the tremendous increase in Internet bandwidth and storage space, video data has been generated, published and spread explosively, becoming an indispensable part of t…
RAG-6DPose: Retrieval-Augmented 6D Pose Estimation via Leveraging CAD as Knowledge Base
Kuanning Wang, Yuqian Fu, Tianyu Wang +4
Accurate 6D pose estimation is key for robotic manipulation, enabling precise object localization for tasks like grasping. We present RAG-6DPose, a retrieval-augmented approach tha…
Rethinking Person Re-identification from a Projection-on-Prototypes Perspective
Qizao Wang, Xuelin Qian, Bin Li +2
Person Re-IDentification (Re-ID) as a retrieval task, has achieved tremendous development over the past decade. Existing state-of-the-art methods follow an analogous framework to f…
Incremental Transformer Structure Enhanced Image Inpainting with Masking Positional Encoding
Qiaole Dong, Chenjie Cao, Yanwei Fu
Image inpainting has made significant advances in recent years. However, it is still challenging to recover corrupted images with both vivid textures and reasonable structures. Som…
HybridGait: A Benchmark for Spatial-Temporal Cloth-Changing Gait Recognition with Hybrid Explorations
Yilan Dong, Chunlin Yu, Ruiyang Ha +5
Existing gait recognition benchmarks mostly include minor clothing variations in the laboratory environments, but lack persistent changes in appearance over time and space. In this…
The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook
Xinlei Yu, Zhangquan Chen, Yongbo He +36
Latent space is rapidly emerging as a native substrate for language-based models. While modern systems are still commonly understood through explicit token-level generation, an inc…
Adaptive Pruning of Pretrained Transformer via Differential Inclusions
Yizhuo Ding, Ke Fan, Yikai Wang +2
Large transformers have demonstrated remarkable success, making it necessary to compress these models to reduce inference costs while preserving their perfor-mance. Current compres…
Doubly Robust Proximal Causal Learning for Continuous Treatments
Yong Wu, Yanwei Fu, Shouyan Wang +1
Proximal causal learning is a promising framework for identifying the causal effect under the existence of unmeasured confounders. Within this framework, the doubly robust (DR) est…
Exploring Efficient Few-shot Adaptation for Vision Transformers
Chengming Xu, Siqian Yang, Yabiao Wang +3
The task of Few-shot Learning (FSL) aims to do the inference on novel categories containing only few labeled examples, with the help of knowledge learned from base categories conta…
Multi-level Semantic Feature Augmentation for One-shot Learning
Zitian Chen, Yanwei Fu, Yinda Zhang +3
The ability to quickly recognize and learn new visual concepts from limited samples enables humans to swiftly adapt to new environments. This ability is enabled by semantic associa…
Robust Classification by Pre-conditioned LASSO and Transductive Diffusion Component Analysis
Yanwei Fu, De-An Huang, Leonid Sigal
Modern machine learning-based recognition approaches require large-scale datasets with large number of labelled training images. However, such datasets are inherently difficult and…
VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation
Sixiao Zheng, Zimian Peng, Yanpeng Zhou +4
Controllable image-to-video (I2V) generation transforms a reference image into a coherent video guided by user-specified control signals. While precise control over camera motion,…
Domain-Aware SE Network for Sketch-based Image Retrieval with Multiplicative Euclidean Margin Softmax
Peng Lu, Gao Huang, Hangyu Lin +3
This paper proposes a novel approach for Sketch-Based Image Retrieval (SBIR), for which the key is to bridge the gap between sketches and photos in terms of the data representation…
Improving Neural Surface Reconstruction with Feature Priors from Multi-View Image
Xinlin Ren, Chenjie Cao, Yanwei Fu +1
Recent advancements in Neural Surface Reconstruction (NSR) have significantly improved multi-view reconstruction when coupled with volume rendering. However, relying solely on phot…
AI Challenger : A Large-scale Dataset for Going Deeper in Image Understanding
Jiahong Wu, He Zheng, Bo Zhao +9
Significant progress has been achieved in Computer Vision by leveraging large-scale image datasets. However, large-scale datasets for complex Computer Vision tasks beyond classific…
Depth Guided Adaptive Meta-Fusion Network for Few-shot Video Recognition
Yuqian Fu, Li Zhang, Junke Wang +2
Humans can easily recognize actions with only a few examples given, while the existing video recognition models still heavily rely on the large-scale labeled data inputs. This obse…
Uni3C: Unifying Precisely 3D-Enhanced Camera and Human Motion Controls for Video Generation
Chenjie Cao, Jingkai Zhou, Shikai Li +5
Camera and human motion controls have been extensively studied for video generation, but existing approaches typically address them separately, suffering from limited data with hig…
VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control
Sixiao Zheng, Minghao Yin, Wenbo Hu +3
Video world models aim to simulate dynamic, real-world environments, yet existing methods struggle to provide unified and precise control over camera and multi-object motion, as vi…
SwiftVideo: A Unified Framework for Few-Step Video Generation through Trajectory-Distribution Alignment
Yanxiao Sun, Jiafu Wu, Yun Cao +6
Diffusion-based or flow-based models have achieved significant progress in video synthesis but require multiple iterative sampling steps, which incurs substantial computational ove…
ReVSeg: Incentivizing the Reasoning Chain for Video Segmentation with Reinforcement Learning
Yifan Li, Yingda Yin, Lingting Zhu +4
Reasoning-centric video object segmentation is an inherently complex task: the query often refers to dynamics, causality, and temporal interactions, rather than static appearances.…
MVSFormer++: Revealing the Devil in Transformer's Details for Multi-View Stereo
Chenjie Cao, Xinlin Ren, Yanwei Fu
Recent advancements in learning-based Multi-View Stereo (MVS) methods have prominently featured transformer-based models with attention mechanisms. However, existing approaches hav…
Robotic Grasping and Placement Controlled by EEG-Based Hybrid Visual and Motor Imagery
Yichang Liu, Tianyu Wang, Ziyi Ye +4
We present a framework that integrates EEG-based visual and motor imagery (VI/MI) with robotic control to enable real-time, intention-driven grasping and placement. Motivated by th…
Joint fMRI Decoding and Encoding with Latent Embedding Alignment
Xuelin Qian, Yikai Wang, Yanwei Fu +3
The connection between brain activity and corresponding visual stimuli is crucial in comprehending the human brain. While deep generative models have exhibited advancement in recov…
Hyper-Transformer for Amodal Completion
Jianxiong Gao, Xuelin Qian, Longfei Liang +2
Amodal object completion is a complex task that involves predicting the invisible parts of an object based on visible segments and background information. Learning shape priors is…
Repositioning the Subject within Image
Yikai Wang, Chenjie Cao, Ke Fan +4
Current image manipulation primarily centers on static manipulation, such as replacing specific regions within an image or altering its overall style. In this paper, we introduce a…
Discovering Causal Relationships using Proxy Variables under Unmeasured Confounding
Yong Wu, Yanwei Fu, Shouyan Wang +2
Inferring causal relationships between variable pairs in the observational study is crucial but challenging, due to the presence of unmeasured confounding. While previous methods e…
Open-DDVM: A Reproduction and Extension of Diffusion Model for Optical Flow Estimation
Qiaole Dong, Bo Zhao, Yanwei Fu
Recently, Google proposes DDVM which for the first time demonstrates that a general diffusion model for image-to-image translation task works impressively well on optical flow esti…
Soft Filter Pruning for Accelerating Deep Convolutional Neural Networks
Yang He, Guoliang Kang, Xuanyi Dong +2
This paper proposed a Soft Filter Pruning (SFP) method to accelerate the inference procedure of deep Convolutional Neural Networks (CNNs). Specifically, the proposed SFP enables th…
Entity-Level Text-Guided Image Manipulation
Yikai Wang, Jianan Wang, Guansong Lu +4
Existing text-guided image manipulation methods aim to modify the appearance of the image or to edit a few objects in a virtual or simple scenario, which is far from practical appl…
MemFlow: Optical Flow Estimation and Prediction with Memory
Qiaole Dong, Yanwei Fu
Optical flow is a classical task that is important to the vision community. Classical optical flow estimation uses two frames as input, whilst some recent methods consider multiple…
PatchMix Augmentation to Identify Causal Features in Few-shot Learning
Chengming Xu, Chen Liu, Xinwei Sun +4
The task of Few-shot learning (FSL) aims to transfer the knowledge learned from base categories with sufficient labelled data to novel categories with scarce known information. It…
V-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence
Jiancheng Pan, Runze Wang, Tianwen Qian +7
Cross-view object correspondence, exemplified by the representative task of ego-exo object correspondence, aims to establish consistent associations of the same object across diffe…
Sequential Multi-Object Grasping with One Dexterous Hand
Sicheng He, Zeyu Shangguan, Kuanning Wang +4
Sequentially grasping multiple objects with multi-fingered hands is common in daily life, where humans can fully leverage the dexterity of their hands to enclose multiple objects.…
NTIRE 2025 Challenge on Cross-Domain Few-Shot Object Detection: Methods and Results
Yuqian Fu, Xingyu Qiu, Bin Ren +59
Cross-Domain Few-Shot Object Detection (CD-FSOD) poses significant challenges to existing object detection and few-shot detection models when applied across domains. In conjunction…
The Blessings of Multiple Treatments and Outcomes in Treatment Effect Estimation
Yong Wu, Mingzhou Liu, Jing Yan +4
Assessing causal effects in the presence of unobserved confounding is a challenging problem. Existing studies leveraged proxy variables or multiple treatments to adjust for the con…
Question Guided Modular Routing Networks for Visual Question Answering
Yanze Wu, Qiang Sun, Jianqi Ma +4
This paper studies the task of Visual Question Answering (VQA), which is topical in Multimedia community recently. Particularly, we explore two critical research problems existed i…
Knockoffs-SPR: Clean Sample Selection in Learning with Noisy Labels
Yikai Wang, Yanwei Fu, Xinwei Sun
A noisy training set usually leads to the degradation of the generalization and robustness of neural networks. In this paper, we propose a novel theoretically guaranteed clean samp…
Vocabulary-informed Zero-shot and Open-set Learning
Yanwei Fu, Xiaomei Wang, Hanze Dong +4
Despite significant progress in object categorization, in recent years, a number of important challenges remain; mainly, the ability to learn from limited labeled data and to recog…
Pushing Auto-regressive Models for 3D Shape Generation at Capacity and Scalability
Xuelin Qian, Yu Wang, Simian Luo +9
Auto-regressive models have achieved impressive results in 2D image generation by modeling joint distributions in grid space. In this paper, we extend auto-regressive models to 3D…
Self-supervised Amodal Video Object Segmentation
Jian Yao, Yuxin Hong, Chiyu Wang +6
Amodal perception requires inferring the full shape of an object that is partially occluded. This task is particularly challenging on two levels: (1) it requires more information t…
DST-Calib: A Dual-Path, Self-Supervised, Target-Free LiDAR-Camera Extrinsic Calibration Network
Zhiwei Huang, Yanwei Fu, Yi Zhou +3
LiDAR-camera extrinsic calibration is essential for multi-modal data fusion in robotic perception systems. However, existing approaches typically rely on handcrafted calibration ta…
Multi-scale Deep Learning Architectures for Person Re-identification
Xuelin Qian, Yanwei Fu, Yu-Gang Jiang +2
Person Re-identification (re-id) aims to match people across non-overlapping camera views in a public space. It is a challenging problem because many people captured in surveillanc…
Online Dense Point Tracking with Streaming Memory
Qiaole Dong, Yanwei Fu
Dense point tracking is a challenging task requiring the continuous tracking of every point in the initial frame throughout a substantial portion of a video, even in the presence o…
DessiLBI: Exploring Structural Sparsity of Deep Networks via Differential Inclusion Paths
Yanwei Fu, Chen Liu, Donghao Li +3
Over-parameterization is ubiquitous nowadays in training neural networks to benefit both optimization in seeking global optima and generalization in reducing prediction error. Howe…
Rapid COVID-19 Risk Screening by Eye-region Manifestations
Yanwei Fu, Lei Zhao, Haojie Zheng +13
It is still nontrivial to develop a new fast COVID-19 screening method with the easier access and lower cost, due to the technical and cost limitations of the current testing metho…
SafeCtrl: Region-Based Safety Control for Text-to-Image Diffusion via Detect-Then-Suppress
Lingyun Zhang, Yu Xie, Yanwei Fu +1
The widespread deployment of text-to-image models is challenged by their potential to generate harmful content. While existing safety methods, such as prompt rewriting or model fin…
LeftRefill: Filling Right Canvas based on Left Reference through Generalized Text-to-Image Diffusion Model
Chenjie Cao, Yunuo Cai, Qiaole Dong +2
This paper introduces LeftRefill, an innovative approach to efficiently harness large Text-to-Image (T2I) diffusion models for reference-guided image synthesis. As the name implies…
LongVie 2: Multimodal Controllable Ultra-Long Video World Model
Jianxiong Gao, Zhaoxi Chen, Xian Liu +7
Building video world models upon pretrained video generation systems represents an important yet challenging step toward general spatiotemporal intelligence. A world model should p…
What Semantics Survive the Connector? Diagnosing VLM-to-DiT Alignment in Video Editing
Hangyu Lin, Chao Wen, Chengming Xu +4
Flow matching based video generative models have been increasingly relying on prepended Vision-Language Models (VLMs) to handle complex, instruction-based video editing. The prevai…
Learning Prior Feature and Attention Enhanced Image Inpainting
Chenjie Cao, Qiaole Dong, Yanwei Fu
Many recent inpainting works have achieved impressive results by leveraging Deep Neural Networks (DNNs) to model various prior information for image restoration. Unfortunately, the…
Schrödinger's Navigator: Imagining an Ensemble of Futures for Zero-Shot Object Navigation
Yu He, Da Huang, Zhenyang Liu +5
Zero-shot object navigation (ZSON) requires robots to find target objects in unseen environments without task-specific fine-tuning or pre-built maps, a key capability for general-p…
ME-D2N: Multi-Expert Domain Decompositional Network for Cross-Domain Few-Shot Learning
Yuqian Fu, Yu Xie, Yanwei Fu +2
Recently, Cross-Domain Few-Shot Learning (CD-FSL) which aims at addressing the Few-Shot Learning (FSL) problem across different domains has attracted rising attention. The core cha…
Cross-Domain Few-Shot Object Detection via Enhanced Open-Set Object Detector
Yuqian Fu, Yu Wang, Yixuan Pan +7
This paper studies the challenging cross-domain few-shot object detection (CD-FSOD), aiming to develop an accurate object detector for novel domains with minimal labeled examples.…
Revisiting Large Language Model Pruning using Neuron Semantic Attribution
Yizhuo Ding, Xinwei Sun, Yanwei Fu +1
Model pruning technique is vital for accelerating large language models by reducing their size and computational requirements. However, the generalizability of existing pruning met…
Domain-RAG: Retrieval-Guided Compositional Image Generation for Cross-Domain Few-Shot Object Detection
Yu Li, Xingyu Qiu, Yuqian Fu +8
Cross-Domain Few-Shot Object Detection (CD-FSOD) aims to detect novel objects with only a handful of labeled samples from previously unseen domains. While data augmentation and gen…
DST: Dynamic Substitute Training for Data-free Black-box Attack
Wenxuan Wang, Xuelin Qian, Yanwei Fu +1
With the wide applications of deep neural network models in various computer vision tasks, more and more works study the model vulnerability to adversarial examples. For data-free…
Local Consensus Enhanced Siamese Network with Reciprocal Loss for Two-view Correspondence Learning
Linbo Wang, Jing Wu, Xianyong Fang +3
Recent studies of two-view correspondence learning usually establish an end-to-end network to jointly predict correspondence reliability and relative pose. We improve such a framew…
MVSFormer: Multi-View Stereo by Learning Robust Image Features and Temperature-based Depth
Chenjie Cao, Xinlin Ren, Yanwei Fu
Feature representation learning is the key recipe for learning-based Multi-View Stereo (MVS). As the common feature extractor of learning-based MVS, vanilla Feature Pyramid Network…
SafeWork-R1: Coevolving Safety and Intelligence under the AI-45 Law
Shanghai AI Lab, :, Yicheng Bao +115
We introduce SafeWork-R1, a cutting-edge multimodal reasoning model that demonstrates the coevolution of capabilities and safety. It is developed by our proposed SafeLadder framewo…
Learning Salient Boundary Feature for Anchor-free Temporal Action Localization
Chuming Lin, Chengming Xu, Donghao Luo +6
Temporal action localization is an important yet challenging task in video understanding. Typically, such a task aims at inferring both the action category and localization of the…
Exploring Structural Sparsity of Deep Networks via Inverse Scale Spaces
Yanwei Fu, Chen Liu, Donghao Li +4
The great success of deep neural networks is built upon their over-parameterization, which smooths the optimization landscape without degrading the generalization ability. Despite…
ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation
Zhenyang Liu, Yongchong Gu, Yikai Wang +2
Recent advances in robot manipulation have leveraged pre-trained vision-language models (VLMs) and explored integrating 3D spatial signals into these models for effective action pr…
Towards Global Optimal Visual In-Context Learning Prompt Selection
Chengming Xu, Chen Liu, Yikai Wang +2
Visual In-Context Learning (VICL) is a prevailing way to transfer visual foundation models to new tasks by leveraging contextual information contained in in-context examples to enh…
Pixel2Mesh++: Multi-View 3D Mesh Generation via Deformation
Chao Wen, Yinda Zhang, Zhuwen Li +1
We study the problem of shape generation in 3D mesh representation from a few color images with known camera poses. While many previous works learn to hallucinate the shape directl…
NI-Tex: Non-isometric Image-based Garment Texture Generation
Hui Shan, Ming Li, Haitao Yang +4
Existing industrial 3D garment meshes already cover most real-world clothing geometries, yet their texture diversity remains limited. To acquire more realistic textures, generative…
VADF: Vision-Adaptive Diffusion Policy Framework for Efficient Robotic Manipulation
Xinglei Yu, Zhenyang Liu, Shufeng Nan +2
Diffusion policies are becoming mainstream in robotic manipulation but suffer from hard negative class imbalance due to uniform sampling and lack of sample difficulty awareness, le…
Learning from Synthetic Data Using a Stacked Multichannel Autoencoder
Xi Zhang, Yanwei Fu, Shanshan Jiang +2
Learning from synthetic data has many important and practical applications. An example of application is photo-sketch recognition. Using synthetic data is challenging due to the di…
Semi-supervised Vocabulary-informed Learning
Yanwei Fu, Leonid Sigal
Despite significant progress in object categorization, in recent years, a number of important challenges remain, mainly, ability to learn from limited labeled data and ability to r…
Grad-PU: Arbitrary-Scale Point Cloud Upsampling via Gradient Descent with Learned Distance Functions
Yun He, Danhang Tang, Yinda Zhang +2
Most existing point cloud upsampling methods have roughly three steps: feature extraction, feature expansion and 3D coordinate prediction. However,they usually suffer from two crit…
Clustering by the Probability Distributions from Extreme Value Theory
Sixiao Zheng, Ke Fan, Yanxi Hou +2
Clustering is an essential task to unsupervised learning. It tries to automatically separate instances into coherent subsets. As one of the most well-known clustering algorithms, k…
Learning to Generate Posters of Scientific Papers
Yuting Qiang, Yanwei Fu, Yanwen Guo +2
Researchers often summarize their work in the form of posters. Posters provide a coherent and efficient way to convey core ideas from scientific papers. Generating a good scientifi…
Vision Transformers: From Semantic Segmentation to Dense Prediction
Li Zhang, Jiachen Lu, Sixiao Zheng +6
The emergence of vision transformers (ViTs) in image classification has shifted the methodologies for visual representation learning. In particular, ViTs learn visual representatio…
Pixel2Mesh++: 3D Mesh Generation and Refinement from Multi-View Images
Chao Wen, Yinda Zhang, Chenjie Cao +3
We study the problem of shape generation in 3D mesh representation from a small number of color images with or without camera poses. While many previous works learn to hallucinate…
Spatial-Temporal Aware Visuomotor Diffusion Policy Learning
Zhenyang Liu, Yikai Wang, Kuanning Wang +3
Visual imitation learning is effective for robots to learn versatile tasks. However, many existing methods rely on behavior cloning with supervised historical trajectories, limitin…
Long-Term Cloth-Changing Person Re-identification
Xuelin Qian, Wenxuan Wang, Li Zhang +5
Person re-identification (Re-ID) aims to match a target person across camera views at different locations and times. Existing Re-ID studies focus on the short-term cloth-consistent…