papers

Publications (262)

cs.CV2020

Neural Pose Transfer by Spatially Adaptive Instance Normalization

Jiashun Wang, Chao Wen, Yanwei Fu +4

Pose transfer has been studied for decades, in which the pose of a source mesh is applied to a target mesh. Particularly in this paper, we are interested in transferring the pose o…

cs.RO2024

Polaris: Open-ended Interactive Robotic Manipulation via Syn2Real Visual Grounding and Large Language Models

Tianyu Wang, Haitao Lin, Junqiu Yu +1

This paper investigates the task of the open-ended interactive robotic manipulation on table-top scenarios. While recent Large Language Models (LLMs) enhance robots' comprehension…

cs.RO2025

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy

Xinyi Chen, Yilun Chen, Yanwei Fu +26

We introduce InternVLA-M1, a unified framework for spatial grounding and robot control that advances instruction-following robots toward scalable, general-purpose intelligence. Its…

cs.CV2025

Towards Reliable and Holistic Visual In-Context Learning Prompt Selection

Wenxiao Wu, Jing-Hao Xue, Chengming Xu +5

Visual In-Context Learning (VICL) has emerged as a prominent approach for adapting visual foundation models to novel tasks, by effectively exploiting contextual information embedde…

cs.CV2017

A Jointly Learned Deep Architecture for Facial Attribute Analysis and Face Detection in the Wild

Keke He, Yanwei Fu, Xiangyang Xue

Facial attribute analysis in the real world scenario is very challenging mainly because of complex face variations. Existing works of analyzing face attributes are mostly based on…

cs.CV2017

Recent Advances in Zero-shot Recognition

Yanwei Fu, Tao Xiang, Yu-Gang Jiang +3

With the recent renaissance of deep convolution neural networks, encouraging breakthroughs have been achieved on the supervised recognition tasks, where each class has sufficient t…

cs.CV2017

Semi-Latent GAN: Learning to generate and modify facial images from attributes

Weidong Yin, Yanwei Fu, Leonid Sigal +1

Generating and manipulating human facial images using high-level attributal controls are important and interesting problems. The models proposed in previous work can solve one of t…

cs.CV2022

ImpDet: Exploring Implicit Fields for 3D Object Detection

Xuelin Qian, Li Wang, Yi Zhu +3

Conventional 3D object detection approaches concentrate on bounding boxes representation learning with several parameters, i.e., localization, dimension, and orientation. Despite i…

cs.CV2019

Image Deformation Meta-Networks for One-Shot Learning

Zitian Chen, Yanwei Fu, Yu-Xiong Wang +3

Humans can robustly learn novel visual concepts even when images undergo various deformations and lose certain information. Mimicking the same behavior and synthesizing deformed in…

cs.RO2026

ST4VLA: Spatially Guided Training for Vision-Language-Action Models

Jinhui Ye, Fangjing Wang, Ning Gao +9

Large vision-language models (VLMs) excel at multimodal understanding but fall short when extended to embodied tasks, where instructions must be transformed into low-level motor ac…

cs.LG2015

Transductive Multi-label Zero-shot Learning

Yanwei Fu, Yongxin Yang, Tim Hospedales +2

Zero-shot learning has received increasing interest as a means to alleviate the often prohibitive expense of annotating training data for large scale recognition problems. These me…

cs.CV2024

FitDiT: Advancing the Authentic Garment Details for High-fidelity Virtual Try-on

Boyuan Jiang, Xiaobin Hu, Donghao Luo +7

Although image-based virtual try-on has made considerable progress, emerging approaches still encounter challenges in producing high-fidelity and robust fitting images across diver…

cs.CV2026

EgoSound: Benchmarking Sound Understanding in Egocentric Videos

Bingwen Zhu, Yuqian Fu, Qiaole Dong +6

Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in vision-language understanding. Yet, human perception is inherently multisensory, integrating…

cs.CV2024

ArtWeaver: Advanced Dynamic Style Integration via Diffusion Model

Chengming Xu, Kai Hu, Qilin Wang +5

Stylized Text-to-Image Generation (STIG) aims to generate images from text prompts and style reference images. In this paper, we present ArtWeaver, a novel framework that leverages…

cs.CV2021

Delving into Data: Effectively Substitute Training for Black-box Attack

Wenxuan Wang, Bangjie Yin, Taiping Yao +6

Deep models have shown their vulnerability when processing adversarial samples. As for the black-box attack, without access to the architecture and weights of the attacked model, t…

cs.CV2018

Stacked Semantic-Guided Attention Model for Fine-Grained Zero-Shot Learning

Yunlong Yu, Zhong Ji, Yanwei Fu +3

Zero-Shot Learning (ZSL) is achieved via aligning the semantic relationships between the global image feature vector and the corresponding class semantic descriptions. However, usi…

cs.CV2024

CustAny: Customizing Anything from A Single Example

Lingjie Kong, Kai Wu, Xiaobin Hu +8

Recent advances in diffusion-based text-to-image models have simplified creating high-fidelity images, but preserving the identity (ID) of specific elements, like a personal dog, i…

cs.CV2025

NSFW-Classifier Guided Prompt Sanitization for Safe Text-to-Image Generation

Yu Xie, Chengjie Zeng, Lingyun Zhang +1

The rapid advancement of text-to-image (T2I) models, such as Stable Diffusion, has enhanced their capability to synthesize images from textual prompts. However, this progress also…

cs.CV2020

Sketch-BERT: Learning Sketch Bidirectional Encoder Representation from Transformers by Self-supervised Learning of Sketch Gestalt

Hangyu Lin, Yanwei Fu, Yu-Gang Jiang +1

Previous researches of sketches often considered sketches in pixel format and leveraged CNN based models in the sketch understanding. Fundamentally, a sketch is stored as a sequenc…

cs.CV2026

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering

Ming Xie, Zizheng Huang, Xudong Tan +6

While streaming omni-video understanding demands continuous perception and proactive, real-time interaction, this crucial area remains largely under-explored. Current omni-modal me…

cs.CV2025

PPMStereo: Pick-and-Play Memory Construction for Consistent Dynamic Stereo Matching

Yun Wang, Junjie Hu, Qiaole Dong +4

Temporally consistent depth estimation from stereo video is critical for real-world applications such as augmented reality, where inconsistent depth estimation disrupts the immersi…

cs.CV2026

Aligned Stable Inpainting: Mitigating Unwanted Object Insertion and Preserving Color Consistency

Yikai Wang, Junqiu Yu, Chenjie Cao +2

Generative image inpainting can produce realistic results even with large, irregular masks, but existing methods still suffer from two common problems: (1) Unwanted object insertio…

cs.CV2022

Local Slot Attention for Vision-and-Language Navigation

Yifeng Zhuang, Qiang Sun, Yanwei Fu +2

Vision-and-language navigation (VLN), a frontier study aiming to pave the way for general-purpose robots, has been a hot topic in the computer vision and natural language processin…

cs.RO2025

TP-MDDN: Task-Preferenced Multi-Demand-Driven Navigation with Autonomous Decision-Making

Shanshan Li, Da Huang, Yu He +3

In daily life, people often move through spaces to find objects that meet their needs, posing a key challenge in embodied AI. Traditional Demand-Driven Navigation (DDN) handles one…

cs.RO2024

SparseGrasp: Robotic Grasping via 3D Semantic Gaussian Splatting from Sparse Multi-View RGB Images

Junqiu Yu, Xinlin Ren, Yongchong Gu +7

Language-guided robotic grasping is a rapidly advancing field where robots are instructed using human language to grasp specific objects. However, existing methods often depend on…

cs.CV2018

Deep Learning for Video Classification and Captioning

Zuxuan Wu, Ting Yao, Yanwei Fu +1

Accelerated by the tremendous increase in Internet bandwidth and storage space, video data has been generated, published and spread explosively, becoming an indispensable part of t…

cs.CV2025

RAG-6DPose: Retrieval-Augmented 6D Pose Estimation via Leveraging CAD as Knowledge Base

Kuanning Wang, Yuqian Fu, Tianyu Wang +4

Accurate 6D pose estimation is key for robotic manipulation, enabling precise object localization for tasks like grasping. We present RAG-6DPose, a retrieval-augmented approach tha…

cs.CV2023

Rethinking Person Re-identification from a Projection-on-Prototypes Perspective

Qizao Wang, Xuelin Qian, Bin Li +2

Person Re-IDentification (Re-ID) as a retrieval task, has achieved tremendous development over the past decade. Existing state-of-the-art methods follow an analogous framework to f…

cs.CV2022

Incremental Transformer Structure Enhanced Image Inpainting with Masking Positional Encoding

Qiaole Dong, Chenjie Cao, Yanwei Fu

Image inpainting has made significant advances in recent years. However, it is still challenging to recover corrupted images with both vivid textures and reasonable structures. Som…

cs.CV2023

HybridGait: A Benchmark for Spatial-Temporal Cloth-Changing Gait Recognition with Hybrid Explorations

Yilan Dong, Chunlin Yu, Ruiyang Ha +5

Existing gait recognition benchmarks mostly include minor clothing variations in the laboratory environments, but lack persistent changes in appearance over time and space. In this…

cs.AI2026

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook

Xinlei Yu, Zhangquan Chen, Yongbo He +36

Latent space is rapidly emerging as a native substrate for language-based models. While modern systems are still commonly understood through explicit token-level generation, an inc…

cs.LG2025

Adaptive Pruning of Pretrained Transformer via Differential Inclusions

Yizhuo Ding, Ke Fan, Yikai Wang +2

Large transformers have demonstrated remarkable success, making it necessary to compress these models to reduce inference costs while preserving their perfor-mance. Current compres…

stat.ME2024

Doubly Robust Proximal Causal Learning for Continuous Treatments

Yong Wu, Yanwei Fu, Shouyan Wang +1

Proximal causal learning is a promising framework for identifying the causal effect under the existence of unmeasured confounders. Within this framework, the doubly robust (DR) est…

cs.CV2023

Exploring Efficient Few-shot Adaptation for Vision Transformers

Chengming Xu, Siqian Yang, Yabiao Wang +3

The task of Few-shot Learning (FSL) aims to do the inference on novel categories containing only few labeled examples, with the help of knowledge learned from base categories conta…

cs.CV2019

Multi-level Semantic Feature Augmentation for One-shot Learning

Zitian Chen, Yanwei Fu, Yinda Zhang +3

The ability to quickly recognize and learn new visual concepts from limited samples enables humans to swiftly adapt to new environments. This ability is enabled by semantic associa…

cs.LG2019

Robust Classification by Pre-conditioned LASSO and Transductive Diffusion Component Analysis

Yanwei Fu, De-An Huang, Leonid Sigal

Modern machine learning-based recognition approaches require large-scale datasets with large number of labelled training images. However, such datasets are inherently difficult and…

cs.CV2026

VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation

Sixiao Zheng, Zimian Peng, Yanpeng Zhou +4

Controllable image-to-video (I2V) generation transforms a reference image into a coherent video guided by user-specified control signals. While precise control over camera motion,…

cs.CV2021

Domain-Aware SE Network for Sketch-based Image Retrieval with Multiplicative Euclidean Margin Softmax

Peng Lu, Gao Huang, Hangyu Lin +3

This paper proposes a novel approach for Sketch-Based Image Retrieval (SBIR), for which the key is to bridge the gap between sketches and photos in terms of the data representation…

cs.CV2024

Improving Neural Surface Reconstruction with Feature Priors from Multi-View Image

Xinlin Ren, Chenjie Cao, Yanwei Fu +1

Recent advancements in Neural Surface Reconstruction (NSR) have significantly improved multi-view reconstruction when coupled with volume rendering. However, relying solely on phot…

cs.CV2017

AI Challenger : A Large-scale Dataset for Going Deeper in Image Understanding

Jiahong Wu, He Zheng, Bo Zhao +9

Significant progress has been achieved in Computer Vision by leveraging large-scale image datasets. However, large-scale datasets for complex Computer Vision tasks beyond classific…

cs.CV2020

Depth Guided Adaptive Meta-Fusion Network for Few-shot Video Recognition

Yuqian Fu, Li Zhang, Junke Wang +2

Humans can easily recognize actions with only a few examples given, while the existing video recognition models still heavily rely on the large-scale labeled data inputs. This obse…

cs.CV2025

Uni3C: Unifying Precisely 3D-Enhanced Camera and Human Motion Controls for Video Generation

Chenjie Cao, Jingkai Zhou, Shikai Li +5

Camera and human motion controls have been extensively studied for video generation, but existing approaches typically address them separately, suffering from limited data with hig…

cs.CV2026

VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control

Sixiao Zheng, Minghao Yin, Wenbo Hu +3

Video world models aim to simulate dynamic, real-world environments, yet existing methods struggle to provide unified and precise control over camera and multi-object motion, as vi…

cs.CV2025

SwiftVideo: A Unified Framework for Few-Step Video Generation through Trajectory-Distribution Alignment

Yanxiao Sun, Jiafu Wu, Yun Cao +6

Diffusion-based or flow-based models have achieved significant progress in video synthesis but require multiple iterative sampling steps, which incurs substantial computational ove…

cs.CV2025

ReVSeg: Incentivizing the Reasoning Chain for Video Segmentation with Reinforcement Learning

Yifan Li, Yingda Yin, Lingting Zhu +4

Reasoning-centric video object segmentation is an inherently complex task: the query often refers to dynamics, causality, and temporal interactions, rather than static appearances.…

cs.CV2024

MVSFormer++: Revealing the Devil in Transformer's Details for Multi-View Stereo

Chenjie Cao, Xinlin Ren, Yanwei Fu

Recent advancements in learning-based Multi-View Stereo (MVS) methods have prominently featured transformer-based models with attention mechanisms. However, existing approaches hav…

cs.RO2026

Robotic Grasping and Placement Controlled by EEG-Based Hybrid Visual and Motor Imagery

Yichang Liu, Tianyu Wang, Ziyi Ye +4

We present a framework that integrates EEG-based visual and motor imagery (VI/MI) with robotic control to enable real-time, intention-driven grasping and placement. Motivated by th…

cs.CV2023

Joint fMRI Decoding and Encoding with Latent Embedding Alignment

Xuelin Qian, Yikai Wang, Yanwei Fu +3

The connection between brain activity and corresponding visual stimuli is crucial in comprehending the human brain. While deep generative models have exhibited advancement in recov…

cs.CV2024

Hyper-Transformer for Amodal Completion

Jianxiong Gao, Xuelin Qian, Longfei Liang +2

Amodal object completion is a complex task that involves predicting the invisible parts of an object based on visible segments and background information. Learning shape priors is…

cs.CV2024

Repositioning the Subject within Image

Yikai Wang, Chenjie Cao, Ke Fan +4

Current image manipulation primarily centers on static manipulation, such as replacing specific regions within an image or altering its overall style. In this paper, we introduce a…

stat.ME2025

Discovering Causal Relationships using Proxy Variables under Unmeasured Confounding

Yong Wu, Yanwei Fu, Shouyan Wang +2

Inferring causal relationships between variable pairs in the observational study is crucial but challenging, due to the presence of unmeasured confounding. While previous methods e…

cs.CV2023

Open-DDVM: A Reproduction and Extension of Diffusion Model for Optical Flow Estimation

Qiaole Dong, Bo Zhao, Yanwei Fu

Recently, Google proposes DDVM which for the first time demonstrates that a general diffusion model for image-to-image translation task works impressively well on optical flow esti…

cs.CV2018

Soft Filter Pruning for Accelerating Deep Convolutional Neural Networks

Yang He, Guoliang Kang, Xuanyi Dong +2

This paper proposed a Soft Filter Pruning (SFP) method to accelerate the inference procedure of deep Convolutional Neural Networks (CNNs). Specifically, the proposed SFP enables th…

cs.CV2023

Entity-Level Text-Guided Image Manipulation

Yikai Wang, Jianan Wang, Guansong Lu +4

Existing text-guided image manipulation methods aim to modify the appearance of the image or to edit a few objects in a virtual or simple scenario, which is far from practical appl…

cs.CV2024

MemFlow: Optical Flow Estimation and Prediction with Memory

Qiaole Dong, Yanwei Fu

Optical flow is a classical task that is important to the vision community. Classical optical flow estimation uses two frames as input, whilst some recent methods consider multiple…

cs.CV2022

PatchMix Augmentation to Identify Causal Features in Few-shot Learning

Chengming Xu, Chen Liu, Xinwei Sun +4

The task of Few-shot learning (FSL) aims to transfer the knowledge learned from base categories with sufficient labelled data to novel categories with scarce known information. It…

cs.CV2026

V-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence

Jiancheng Pan, Runze Wang, Tianwen Qian +7

Cross-view object correspondence, exemplified by the representative task of ego-exo object correspondence, aims to establish consistent associations of the same object across diffe…

cs.RO2025

Sequential Multi-Object Grasping with One Dexterous Hand

Sicheng He, Zeyu Shangguan, Kuanning Wang +4

Sequentially grasping multiple objects with multi-fingered hands is common in daily life, where humans can fully leverage the dexterity of their hands to enclose multiple objects.…

cs.CV2025

NTIRE 2025 Challenge on Cross-Domain Few-Shot Object Detection: Methods and Results

Yuqian Fu, Xingyu Qiu, Bin Ren +59

Cross-Domain Few-Shot Object Detection (CD-FSOD) poses significant challenges to existing object detection and few-shot detection models when applied across domains. In conjunction…

stat.ME2023

The Blessings of Multiple Treatments and Outcomes in Treatment Effect Estimation

Yong Wu, Mingzhou Liu, Jing Yan +4

Assessing causal effects in the presence of unobserved confounding is a challenging problem. Existing studies leveraged proxy variables or multiple treatments to adjust for the con…

cs.CV2020

Question Guided Modular Routing Networks for Visual Question Answering

Yanze Wu, Qiang Sun, Jianqi Ma +4

This paper studies the task of Visual Question Answering (VQA), which is topical in Multimedia community recently. Particularly, we explore two critical research problems existed i…

cs.LG2023

Knockoffs-SPR: Clean Sample Selection in Learning with Noisy Labels

Yikai Wang, Yanwei Fu, Xinwei Sun

A noisy training set usually leads to the degradation of the generalization and robustness of neural networks. In this paper, we propose a novel theoretically guaranteed clean samp…

cs.CV2023

Vocabulary-informed Zero-shot and Open-set Learning

Yanwei Fu, Xiaomei Wang, Hanze Dong +4

Despite significant progress in object categorization, in recent years, a number of important challenges remain; mainly, the ability to learn from limited labeled data and to recog…

cs.CV2024

Pushing Auto-regressive Models for 3D Shape Generation at Capacity and Scalability

Xuelin Qian, Yu Wang, Simian Luo +9

Auto-regressive models have achieved impressive results in 2D image generation by modeling joint distributions in grid space. In this paper, we extend auto-regressive models to 3D…

cs.CV2022

Self-supervised Amodal Video Object Segmentation

Jian Yao, Yuxin Hong, Chiyu Wang +6

Amodal perception requires inferring the full shape of an object that is partially occluded. This task is particularly challenging on two levels: (1) it requires more information t…

cs.RO2026

DST-Calib: A Dual-Path, Self-Supervised, Target-Free LiDAR-Camera Extrinsic Calibration Network

Zhiwei Huang, Yanwei Fu, Yi Zhou +3

LiDAR-camera extrinsic calibration is essential for multi-modal data fusion in robotic perception systems. However, existing approaches typically rely on handcrafted calibration ta…

cs.CV2017

Multi-scale Deep Learning Architectures for Person Re-identification

Xuelin Qian, Yanwei Fu, Yu-Gang Jiang +2

Person Re-identification (re-id) aims to match people across non-overlapping camera views in a public space. It is a challenging problem because many people captured in surveillanc…

cs.CV2025

Online Dense Point Tracking with Streaming Memory

Qiaole Dong, Yanwei Fu

Dense point tracking is a challenging task requiring the continuous tracking of every point in the initial frame throughout a substantial portion of a video, even in the presence o…

cs.CV2020

DessiLBI: Exploring Structural Sparsity of Deep Networks via Differential Inclusion Paths

Yanwei Fu, Chen Liu, Donghao Li +3

Over-parameterization is ubiquitous nowadays in training neural networks to benefit both optimization in seeking global optima and generalization in reducing prediction error. Howe…

eess.IV2021

Rapid COVID-19 Risk Screening by Eye-region Manifestations

Yanwei Fu, Lei Zhao, Haojie Zheng +13

It is still nontrivial to develop a new fast COVID-19 screening method with the easier access and lower cost, due to the technical and cost limitations of the current testing metho…

cs.CV2025

SafeCtrl: Region-Based Safety Control for Text-to-Image Diffusion via Detect-Then-Suppress

Lingyun Zhang, Yu Xie, Yanwei Fu +1

The widespread deployment of text-to-image models is challenged by their potential to generate harmful content. While existing safety methods, such as prompt rewriting or model fin…

cs.CV2024

LeftRefill: Filling Right Canvas based on Left Reference through Generalized Text-to-Image Diffusion Model

Chenjie Cao, Yunuo Cai, Qiaole Dong +2

This paper introduces LeftRefill, an innovative approach to efficiently harness large Text-to-Image (T2I) diffusion models for reference-guided image synthesis. As the name implies…

cs.CV2025

LongVie 2: Multimodal Controllable Ultra-Long Video World Model

Jianxiong Gao, Zhaoxi Chen, Xian Liu +7

Building video world models upon pretrained video generation systems represents an important yet challenging step toward general spatiotemporal intelligence. A world model should p…

cs.CV2026

What Semantics Survive the Connector? Diagnosing VLM-to-DiT Alignment in Video Editing

Hangyu Lin, Chao Wen, Chengming Xu +4

Flow matching based video generative models have been increasingly relying on prepended Vision-Language Models (VLMs) to handle complex, instruction-based video editing. The prevai…

cs.CV2023

Learning Prior Feature and Attention Enhanced Image Inpainting

Chenjie Cao, Qiaole Dong, Yanwei Fu

Many recent inpainting works have achieved impressive results by leveraging Deep Neural Networks (DNNs) to model various prior information for image restoration. Unfortunately, the…

cs.RO2026

Schrödinger's Navigator: Imagining an Ensemble of Futures for Zero-Shot Object Navigation

Yu He, Da Huang, Zhenyang Liu +5

Zero-shot object navigation (ZSON) requires robots to find target objects in unseen environments without task-specific fine-tuning or pre-built maps, a key capability for general-p…

cs.CV2022

ME-D2N: Multi-Expert Domain Decompositional Network for Cross-Domain Few-Shot Learning

Yuqian Fu, Yu Xie, Yanwei Fu +2

Recently, Cross-Domain Few-Shot Learning (CD-FSL) which aims at addressing the Few-Shot Learning (FSL) problem across different domains has attracted rising attention. The core cha…

cs.CV2024

Cross-Domain Few-Shot Object Detection via Enhanced Open-Set Object Detector

Yuqian Fu, Yu Wang, Yixuan Pan +7

This paper studies the challenging cross-domain few-shot object detection (CD-FSOD), aiming to develop an accurate object detector for novel domains with minimal labeled examples.…

cs.CL2025

Revisiting Large Language Model Pruning using Neuron Semantic Attribution

Yizhuo Ding, Xinwei Sun, Yanwei Fu +1

Model pruning technique is vital for accelerating large language models by reducing their size and computational requirements. However, the generalizability of existing pruning met…

cs.CV2025

Domain-RAG: Retrieval-Guided Compositional Image Generation for Cross-Domain Few-Shot Object Detection

Yu Li, Xingyu Qiu, Yuqian Fu +8

Cross-Domain Few-Shot Object Detection (CD-FSOD) aims to detect novel objects with only a handful of labeled samples from previously unseen domains. While data augmentation and gen…

cs.CV2022

DST: Dynamic Substitute Training for Data-free Black-box Attack

Wenxuan Wang, Xuelin Qian, Yanwei Fu +1

With the wide applications of deep neural network models in various computer vision tasks, more and more works study the model vulnerability to adversarial examples. For data-free…

cs.CV2023

Local Consensus Enhanced Siamese Network with Reciprocal Loss for Two-view Correspondence Learning

Linbo Wang, Jing Wu, Xianyong Fang +3

Recent studies of two-view correspondence learning usually establish an end-to-end network to jointly predict correspondence reliability and relative pose. We improve such a framew…

cs.CV2022

MVSFormer: Multi-View Stereo by Learning Robust Image Features and Temperature-based Depth

Chenjie Cao, Xinlin Ren, Yanwei Fu

Feature representation learning is the key recipe for learning-based Multi-View Stereo (MVS). As the common feature extractor of learning-based MVS, vanilla Feature Pyramid Network…

cs.AI2025

SafeWork-R1: Coevolving Safety and Intelligence under the AI-45 Law

Shanghai AI Lab, :, Yicheng Bao +115

We introduce SafeWork-R1, a cutting-edge multimodal reasoning model that demonstrates the coevolution of capabilities and safety. It is developed by our proposed SafeLadder framewo…

cs.CV2021

Learning Salient Boundary Feature for Anchor-free Temporal Action Localization

Chuming Lin, Chengming Xu, Donghao Luo +6

Temporal action localization is an important yet challenging task in video understanding. Typically, such a task aims at inferring both the action category and localization of the…

cs.LG2022

Exploring Structural Sparsity of Deep Networks via Inverse Scale Spaces

Yanwei Fu, Chen Liu, Donghao Li +4

The great success of deep neural networks is built upon their over-parameterization, which smooths the optimization landscape without degrading the generalization ability. Despite…

cs.RO2026

ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation

Zhenyang Liu, Yongchong Gu, Yikai Wang +2

Recent advances in robot manipulation have leveraged pre-trained vision-language models (VLMs) and explored integrating 3D spatial signals into these models for effective action pr…

cs.CV2024

Towards Global Optimal Visual In-Context Learning Prompt Selection

Chengming Xu, Chen Liu, Yikai Wang +2

Visual In-Context Learning (VICL) is a prevailing way to transfer visual foundation models to new tasks by leveraging contextual information contained in in-context examples to enh…

cs.CV2019

Pixel2Mesh++: Multi-View 3D Mesh Generation via Deformation

Chao Wen, Yinda Zhang, Zhuwen Li +1

We study the problem of shape generation in 3D mesh representation from a few color images with known camera poses. While many previous works learn to hallucinate the shape directl…

cs.CV2026

NI-Tex: Non-isometric Image-based Garment Texture Generation

Hui Shan, Ming Li, Haitao Yang +4

Existing industrial 3D garment meshes already cover most real-world clothing geometries, yet their texture diversity remains limited. To acquire more realistic textures, generative…

cs.RO2026

VADF: Vision-Adaptive Diffusion Policy Framework for Efficient Robotic Manipulation

Xinglei Yu, Zhenyang Liu, Shufeng Nan +2

Diffusion policies are becoming mainstream in robotic manipulation but suffer from hard negative class imbalance due to uniform sampling and lack of sample difficulty awareness, le…

cs.CV2015

Learning from Synthetic Data Using a Stacked Multichannel Autoencoder

Xi Zhang, Yanwei Fu, Shanshan Jiang +2

Learning from synthetic data has many important and practical applications. An example of application is photo-sketch recognition. Using synthetic data is challenging due to the di…

cs.CV2016

Semi-supervised Vocabulary-informed Learning

Yanwei Fu, Leonid Sigal

Despite significant progress in object categorization, in recent years, a number of important challenges remain, mainly, ability to learn from limited labeled data and ability to r…

cs.CV2023

Grad-PU: Arbitrary-Scale Point Cloud Upsampling via Gradient Descent with Learned Distance Functions

Yun He, Danhang Tang, Yinda Zhang +2

Most existing point cloud upsampling methods have roughly three steps: feature extraction, feature expansion and 3D coordinate prediction. However,they usually suffer from two crit…

cs.LG2022

Clustering by the Probability Distributions from Extreme Value Theory

Sixiao Zheng, Ke Fan, Yanxi Hou +2

Clustering is an essential task to unsupervised learning. It tries to automatically separate instances into coherent subsets. As one of the most well-known clustering algorithms, k…

cs.AI2016

Learning to Generate Posters of Scientific Papers

Yuting Qiang, Yanwei Fu, Yanwen Guo +2

Researchers often summarize their work in the form of posters. Posters provide a coherent and efficient way to convey core ideas from scientific papers. Generating a good scientifi…

cs.CV2024

Vision Transformers: From Semantic Segmentation to Dense Prediction

Li Zhang, Jiachen Lu, Sixiao Zheng +6

The emergence of vision transformers (ViTs) in image classification has shifted the methodologies for visual representation learning. In particular, ViTs learn visual representatio…

cs.CV2022

Pixel2Mesh++: 3D Mesh Generation and Refinement from Multi-View Images

Chao Wen, Yinda Zhang, Chenjie Cao +3

We study the problem of shape generation in 3D mesh representation from a small number of color images with or without camera poses. While many previous works learn to hallucinate…

cs.RO2025

Spatial-Temporal Aware Visuomotor Diffusion Policy Learning

Zhenyang Liu, Yikai Wang, Kuanning Wang +3

Visual imitation learning is effective for robots to learn versatile tasks. However, many existing methods rely on behavior cloning with supervised historical trajectories, limitin…

cs.CV2020

Long-Term Cloth-Changing Person Re-identification

Xuelin Qian, Wenxuan Wang, Li Zhang +5

Person re-identification (Re-ID) aims to match a target person across camera views at different locations and times. Existing Re-ID studies focus on the short-term cloth-consistent…