papers

Publications (40)

cs.CV2024

SC4D: Sparse-Controlled Video-to-4D Generation and Motion Transfer

Zijie Wu, Chaohui Yu, Yanqin Jiang +3

Recent advances in 2D/3D generative models enable the generation of dynamic 3D objects from a single-view video. Existing approaches utilize score distillation sampling to form the…

cs.CV2023

SingleInsert: Inserting New Concepts from a Single Image into Text-to-Image Models for Flexible Editing

Zijie Wu, Chaohui Yu, Zhen Zhu +2

Recent progress in text-to-image (T2I) models enables high-quality image generation with flexible textual control. To utilize the abundant visual priors in the off-the-shelf T2I mo…

cs.CV2026

Astra: a generalizable report generation foundation model for 3D computed tomography

Zhuhao Wang, Fang Chen, Chaohui Yu +19

Astra is a foundation model that automatically generates radiology reports from thoracoabdominal CT scans, handling multiple organ regions and maintaining consistent style across d…

#ct report generation#medical imaging#foundation models#reinforcement learning
cs.CV2023

RegionBLIP: A Unified Multi-modal Pre-training Framework for Holistic and Regional Comprehension

Qiang Zhou, Chaohui Yu, Shaofeng Zhang +3

In this work, we investigate extending the comprehension of Multi-modal Large Language Models (MLLMs) to regional objects. To this end, we propose to extract features corresponding…

cs.CV2023

Foundation Model Drives Weakly Incremental Learning for Semantic Segmentation

Chaohui Yu, Qiang Zhou, Jingliang Li +3

Modern incremental learning for semantic segmentation methods usually learn new categories based on dense annotations. Although achieve promising results, pixel-by-pixel labeling i…

cs.LG2019

Accelerating Deep Unsupervised Domain Adaptation with Transfer Channel Pruning

Chaohui Yu, Jindong Wang, Yiqiang Chen +1

Deep unsupervised domain adaptation (UDA) has recently received increasing attention from researchers. However, existing methods are computationally intensive due to the computatio…

cs.CV2023

MimCo: Masked Image Modeling Pre-training with Contrastive Teacher

Qiang Zhou, Chaohui Yu, Hao Luo +2

Recent masked image modeling (MIM) has received much attention in self-supervised learning (SSL), which requires the target model to recover the masked part of the input image. Alt…

cs.LG2020

Learning to Match Distributions for Domain Adaptation

Chaohui Yu, Jindong Wang, Chang Liu +5

When the training and test data are from different distributions, domain adaptation is needed to reduce dataset bias to improve the model's generalization ability. Since it is diff…

cs.CV2024

MVInpainter: Learning Multi-View Consistent Inpainting to Bridge 2D and 3D Editing

Chenjie Cao, Chaohui Yu, Fan Wang +2

Novel View Synthesis (NVS) and 3D generation have recently achieved prominent improvements. However, these works mainly focus on confined categories or synthetic 3D assets, which a…

cs.LG2021

FedHealth: A Federated Transfer Learning Framework for Wearable Healthcare

Yiqiang Chen, Jindong Wang, Chaohui Yu +2

With the rapid development of computing technology, wearable devices such as smart phones and wristbands make it easy to get access to people's health information including activit…

cs.CV2025

MVGenMaster: Scaling Multi-View Generation from Any Image via 3D Priors Enhanced Diffusion Model

Chenjie Cao, Chaohui Yu, Shang Liu +3

We introduce MVGenMaster, a multi-view diffusion model enhanced with 3D priors to address versatile Novel View Synthesis (NVS) tasks. MVGenMaster leverages 3D priors that are warpe…

cs.CV2024

Animate3D: Animating Any 3D Model with Multi-view Video Diffusion

Yanqin Jiang, Chaohui Yu, Chenjie Cao +3

Recent advances in 4D generation mainly focus on generating 4D content by distilling pre-trained text or single-view image-conditioned models. It is inconvenient for them to take a…

cs.CV2026

AnimateAnyMesh++: A Flexible 4D Foundation Model for High-Fidelity Text-Driven Mesh Animation

Zijie Wu, Chaohui Yu, Fan Wang +1

Recent advances in 4D content generation have attracted increasing attention, yet creating high-quality animated 3D models remains challenging due to the complexity of modeling spa…

cs.CV2023

ICPC: Instance-Conditioned Prompting with Contrastive Learning for Semantic Segmentation

Chaohui Yu, Qiang Zhou, Zhibin Wang +1

Modern supervised semantic segmentation methods are usually finetuned based on the supervised or self-supervised models pre-trained on ImageNet. Recent work shows that transferring…

cs.CV2021

Instant-Teaching: An End-to-End Semi-Supervised Object Detection Framework

Qiang Zhou, Chaohui Yu, Zhibin Wang +2

Supervised learning based object detection frameworks demand plenty of laborious manual annotations, which may not be practical in real applications. Semi-supervised object detecti…

cs.CV2024

MeshSegmenter: Zero-Shot Mesh Semantic Segmentation via Texture Synthesis

Ziming Zhong, Yanxu Xu, Jing Li +4

We present MeshSegmenter, a simple yet effective framework designed for zero-shot 3D semantic segmentation. This model successfully extends the powerful capabilities of 2D segmenta…

cs.CV2026

Vascular anatomy-aware self-supervised pre-training for X-ray angiogram analysis

De-Xing Huang, Chaohui Yu, Xiao-Hu Zhou +8

X-ray angiography is the gold standard imaging modality for cardiovascular diseases. However, current deep learning approaches for X-ray angiogram analysis are severely constrained…

cs.CV2025

3DV-TON: Textured 3D-Guided Consistent Video Try-on via Diffusion Models

Min Wei, Chaohui Yu, Jingkai Zhou +1

Video try-on replaces clothing in videos with target garments. Existing methods struggle to generate high-quality and temporally consistent results when handling complex clothing p…

cs.CV2023

ES-MVSNet: Efficient Framework for End-to-end Self-supervised Multi-View Stereo

Qiang Zhou, Chaohui Yu, Jingliang Li +3

Compared to the multi-stage self-supervised multi-view stereo (MVS) method, the end-to-end (E2E) approach has received more attention due to its concise and efficient training pipe…

cs.CV2025

LiON-LoRA: Rethinking LoRA Fusion to Unify Controllable Spatial and Temporal Generation for Video Diffusion

Yisu Zhang, Chenjie Cao, Chaohui Yu +1

Video Diffusion Models (VDMs) have demonstrated remarkable capabilities in synthesizing realistic videos by learning from large-scale data. Although vanilla Low-Rank Adaptation (Lo…

cs.LG2019

Transfer Learning with Dynamic Adversarial Adaptation Network

Chaohui Yu, Jindong Wang, Yiqiang Chen +1

The recent advances in deep transfer learning reveal that adversarial learning can be embedded into deep networks to learn more transferable features to reduce the distribution dis…

cs.CV2025

AnimateAnyMesh: A Feed-Forward 4D Foundation Model for Text-Driven Universal Mesh Animation

Zijie Wu, Chaohui Yu, Fan Wang +1

Recent advances in 4D content generation have attracted increasing attention, yet creating high-quality animated 3D models remains challenging due to the complexity of modeling spa…

cs.CV2022

Point RCNN: An Angle-Free Framework for Rotated Object Detection

Qiang Zhou, Chaohui Yu, Zhibin Wang +1

Rotated object detection in aerial images is still challenging due to arbitrary orientations, large scale and aspect ratio variations, and extreme density of objects. Existing stat…

cs.CV2023

D2Q-DETR: Decoupling and Dynamic Queries for Oriented Object Detection with Transformers

Qiang Zhou, Chaohui Yu, Zhibin Wang +1

Despite the promising results, existing oriented object detection methods usually involve heuristically designed rules, e.g., RRoI generation, rotated NMS. In this paper, we propos…

cs.CV2025

Cyc3D: Fine-grained Controllable 3D Generation via Cycle Consistency Regularization

Hongbin Xu, Chaohui Yu, Feng Xiao +5

Despite the remarkable progress of 3D generation, achieving controllability, i.e., ensuring consistency between generated 3D content and input conditions like edge and depth, remai…

cs.CV2023

LMSeg: Language-guided Multi-dataset Segmentation

Qiang Zhou, Yuang Liu, Chaohui Yu +3

It's a meaningful and attractive topic to build a general and inclusive segmentation model that can recognize more categories in various scenarios. A straightforward way is to comb…

cs.CV2021

Object Detection Made Simpler by Eliminating Heuristic NMS

Qiang Zhou, Chaohui Yu, Chunhua Shen +2

We show a simple NMS-free, end-to-end object detection framework, of which the network is a minimal modification to a one-stage object detector such as the FCOS detection model [Ti…

cs.CV2025

EarthCrafter: Scalable 3D Earth Generation via Dual-Sparse Latent Diffusion

Shang Liu, Chenjie Cao, Chaohui Yu +3

Despite the remarkable developments achieved by recent 3D generation works, scaling these methods to geographic extents, such as modeling thousands of square kilometers of Earth's…

cs.CV2025

GeoVideo: Introducing Geometric Regularization into Video Generation Model

Yunpeng Bai, Shaoheng Fang, Chaohui Yu +2

Recent advances in video generation have enabled the synthesis of high-quality and visually realistic clips using diffusion transformer models. However, most existing approaches op…

eess.IV2021

Learning Invariant Representations across Domains and Tasks

Jindong Wang, Wenjie Feng, Chang Liu +5

Being expensive and time-consuming to collect massive COVID-19 image samples to train deep classification models, transfer learning is a promising approach by transferring knowledg…

cs.RO2026

RynnVLA-002: A Unified Vision-Language-Action and World Model

Jun Cen, Siteng Huang, Yuqian Yuan +11

We introduce RynnVLA-002, a unified Vision-Language-Action (VLA) and world model. The world model leverages action and visual inputs to predict future image states, learning the un…

cs.CV2026

CamPVG: Camera-Controlled Panoramic Video Generation with Epipolar-Aware Diffusion

Chenhao Ji, Chaohui Yu, Junyao Gao +2

Recently, camera-controlled video generation has seen rapid development, offering more precise control over video generation. However, existing methods predominantly focus on camer…

cs.CV2025

MVRoom: Controllable 3D Indoor Scene Generation with Multi-View Diffusion Models

Shaoheng Fang, Chaohui Yu, Fan Wang +1

We introduce MVRoom, a controllable novel view synthesis (NVS) pipeline for 3D indoor scenes that uses multi-view diffusion conditioned on a coarse 3D layout. MVRoom employs a two-…

cs.CV2023

Improved Neural Radiance Fields Using Pseudo-depth and Fusion

Jingliang Li, Qiang Zhou, Chaohui Yu +4

Since the advent of Neural Radiance Fields, novel view synthesis has received tremendous attention. The existing approach for the generalization of radiance field reconstruction pr…

cs.CV2026

SCOPE: Scale-Consistent One-Pass Estimation of 3D Geometry

Zheng Zhang, Lihe Yang, Tianyu Yang +6

We present SCOPE (Scale-Consistent One-Pass Estimation of 3D Geometry), a novel approach for estimating 3D geometry from extended monocular video sequences, where existing methods…

cs.RO2025

WorldVLA: Towards Autoregressive Action World Model

Jun Cen, Chaohui Yu, Hangjie Yuan +9

We present WorldVLA, an autoregressive action world model that unifies action and image understanding and generation. Our WorldVLA intergrates Vision-Language-Action (VLA) model an…

cs.CV2024

VCD-Texture: Variance Alignment based 3D-2D Co-Denoising for Text-Guided Texturing

Shang Liu, Chaohui Yu, Chenjie Cao +2

Recent research on texture synthesis for 3D shapes benefits a lot from dramatically developed 2D text-to-image diffusion models, including inpainting-based and optimization-based a…

cs.CV2023

Points-to-3D: Bridging the Gap between Sparse Points and Shape-Controllable Text-to-3D Generation

Chaohui Yu, Qiang Zhou, Jingliang Li +3

Text-to-3D generation has recently garnered significant attention, fueled by 2D diffusion models trained on billions of image-text pairs. Existing methods primarily rely on score d…

cs.CV2026

A Contrastive Pre-trained Foundation Model for Deciphering Imaging Noisomics across Modalities

Yuanjie Gu, Yiqun Wang, Chaohui Yu +4

Characterizing imaging noise is notoriously data-intensive and device-dependent, as modern sensors entangle physical signals with complex algorithmic artifacts. Current paradigms s…

cs.CV2025

Uni3C: Unifying Precisely 3D-Enhanced Camera and Human Motion Controls for Video Generation

Chenjie Cao, Jingkai Zhou, Shikai Li +5

Camera and human motion controls have been extensively studied for video generation, but existing approaches typically address them separately, suffering from limited data with hig…