Publications (40)
SC4D: Sparse-Controlled Video-to-4D Generation and Motion Transfer
Zijie Wu, Chaohui Yu, Yanqin Jiang +3
Recent advances in 2D/3D generative models enable the generation of dynamic 3D objects from a single-view video. Existing approaches utilize score distillation sampling to form the…
SingleInsert: Inserting New Concepts from a Single Image into Text-to-Image Models for Flexible Editing
Zijie Wu, Chaohui Yu, Zhen Zhu +2
Recent progress in text-to-image (T2I) models enables high-quality image generation with flexible textual control. To utilize the abundant visual priors in the off-the-shelf T2I mo…
Astra: a generalizable report generation foundation model for 3D computed tomography
Zhuhao Wang, Fang Chen, Chaohui Yu +19
Astra is a foundation model that automatically generates radiology reports from thoracoabdominal CT scans, handling multiple organ regions and maintaining consistent style across d…
RegionBLIP: A Unified Multi-modal Pre-training Framework for Holistic and Regional Comprehension
Qiang Zhou, Chaohui Yu, Shaofeng Zhang +3
In this work, we investigate extending the comprehension of Multi-modal Large Language Models (MLLMs) to regional objects. To this end, we propose to extract features corresponding…
Foundation Model Drives Weakly Incremental Learning for Semantic Segmentation
Chaohui Yu, Qiang Zhou, Jingliang Li +3
Modern incremental learning for semantic segmentation methods usually learn new categories based on dense annotations. Although achieve promising results, pixel-by-pixel labeling i…
Accelerating Deep Unsupervised Domain Adaptation with Transfer Channel Pruning
Chaohui Yu, Jindong Wang, Yiqiang Chen +1
Deep unsupervised domain adaptation (UDA) has recently received increasing attention from researchers. However, existing methods are computationally intensive due to the computatio…
MimCo: Masked Image Modeling Pre-training with Contrastive Teacher
Qiang Zhou, Chaohui Yu, Hao Luo +2
Recent masked image modeling (MIM) has received much attention in self-supervised learning (SSL), which requires the target model to recover the masked part of the input image. Alt…
Learning to Match Distributions for Domain Adaptation
Chaohui Yu, Jindong Wang, Chang Liu +5
When the training and test data are from different distributions, domain adaptation is needed to reduce dataset bias to improve the model's generalization ability. Since it is diff…
MVInpainter: Learning Multi-View Consistent Inpainting to Bridge 2D and 3D Editing
Chenjie Cao, Chaohui Yu, Fan Wang +2
Novel View Synthesis (NVS) and 3D generation have recently achieved prominent improvements. However, these works mainly focus on confined categories or synthetic 3D assets, which a…
FedHealth: A Federated Transfer Learning Framework for Wearable Healthcare
Yiqiang Chen, Jindong Wang, Chaohui Yu +2
With the rapid development of computing technology, wearable devices such as smart phones and wristbands make it easy to get access to people's health information including activit…
MVGenMaster: Scaling Multi-View Generation from Any Image via 3D Priors Enhanced Diffusion Model
Chenjie Cao, Chaohui Yu, Shang Liu +3
We introduce MVGenMaster, a multi-view diffusion model enhanced with 3D priors to address versatile Novel View Synthesis (NVS) tasks. MVGenMaster leverages 3D priors that are warpe…
Animate3D: Animating Any 3D Model with Multi-view Video Diffusion
Yanqin Jiang, Chaohui Yu, Chenjie Cao +3
Recent advances in 4D generation mainly focus on generating 4D content by distilling pre-trained text or single-view image-conditioned models. It is inconvenient for them to take a…
AnimateAnyMesh++: A Flexible 4D Foundation Model for High-Fidelity Text-Driven Mesh Animation
Zijie Wu, Chaohui Yu, Fan Wang +1
Recent advances in 4D content generation have attracted increasing attention, yet creating high-quality animated 3D models remains challenging due to the complexity of modeling spa…
ICPC: Instance-Conditioned Prompting with Contrastive Learning for Semantic Segmentation
Chaohui Yu, Qiang Zhou, Zhibin Wang +1
Modern supervised semantic segmentation methods are usually finetuned based on the supervised or self-supervised models pre-trained on ImageNet. Recent work shows that transferring…
Instant-Teaching: An End-to-End Semi-Supervised Object Detection Framework
Qiang Zhou, Chaohui Yu, Zhibin Wang +2
Supervised learning based object detection frameworks demand plenty of laborious manual annotations, which may not be practical in real applications. Semi-supervised object detecti…
MeshSegmenter: Zero-Shot Mesh Semantic Segmentation via Texture Synthesis
Ziming Zhong, Yanxu Xu, Jing Li +4
We present MeshSegmenter, a simple yet effective framework designed for zero-shot 3D semantic segmentation. This model successfully extends the powerful capabilities of 2D segmenta…
Vascular anatomy-aware self-supervised pre-training for X-ray angiogram analysis
De-Xing Huang, Chaohui Yu, Xiao-Hu Zhou +8
X-ray angiography is the gold standard imaging modality for cardiovascular diseases. However, current deep learning approaches for X-ray angiogram analysis are severely constrained…
3DV-TON: Textured 3D-Guided Consistent Video Try-on via Diffusion Models
Min Wei, Chaohui Yu, Jingkai Zhou +1
Video try-on replaces clothing in videos with target garments. Existing methods struggle to generate high-quality and temporally consistent results when handling complex clothing p…
ES-MVSNet: Efficient Framework for End-to-end Self-supervised Multi-View Stereo
Qiang Zhou, Chaohui Yu, Jingliang Li +3
Compared to the multi-stage self-supervised multi-view stereo (MVS) method, the end-to-end (E2E) approach has received more attention due to its concise and efficient training pipe…
LiON-LoRA: Rethinking LoRA Fusion to Unify Controllable Spatial and Temporal Generation for Video Diffusion
Yisu Zhang, Chenjie Cao, Chaohui Yu +1
Video Diffusion Models (VDMs) have demonstrated remarkable capabilities in synthesizing realistic videos by learning from large-scale data. Although vanilla Low-Rank Adaptation (Lo…
Transfer Learning with Dynamic Adversarial Adaptation Network
Chaohui Yu, Jindong Wang, Yiqiang Chen +1
The recent advances in deep transfer learning reveal that adversarial learning can be embedded into deep networks to learn more transferable features to reduce the distribution dis…
AnimateAnyMesh: A Feed-Forward 4D Foundation Model for Text-Driven Universal Mesh Animation
Zijie Wu, Chaohui Yu, Fan Wang +1
Recent advances in 4D content generation have attracted increasing attention, yet creating high-quality animated 3D models remains challenging due to the complexity of modeling spa…
Point RCNN: An Angle-Free Framework for Rotated Object Detection
Qiang Zhou, Chaohui Yu, Zhibin Wang +1
Rotated object detection in aerial images is still challenging due to arbitrary orientations, large scale and aspect ratio variations, and extreme density of objects. Existing stat…
D2Q-DETR: Decoupling and Dynamic Queries for Oriented Object Detection with Transformers
Qiang Zhou, Chaohui Yu, Zhibin Wang +1
Despite the promising results, existing oriented object detection methods usually involve heuristically designed rules, e.g., RRoI generation, rotated NMS. In this paper, we propos…
Cyc3D: Fine-grained Controllable 3D Generation via Cycle Consistency Regularization
Hongbin Xu, Chaohui Yu, Feng Xiao +5
Despite the remarkable progress of 3D generation, achieving controllability, i.e., ensuring consistency between generated 3D content and input conditions like edge and depth, remai…
LMSeg: Language-guided Multi-dataset Segmentation
Qiang Zhou, Yuang Liu, Chaohui Yu +3
It's a meaningful and attractive topic to build a general and inclusive segmentation model that can recognize more categories in various scenarios. A straightforward way is to comb…
Object Detection Made Simpler by Eliminating Heuristic NMS
Qiang Zhou, Chaohui Yu, Chunhua Shen +2
We show a simple NMS-free, end-to-end object detection framework, of which the network is a minimal modification to a one-stage object detector such as the FCOS detection model [Ti…
EarthCrafter: Scalable 3D Earth Generation via Dual-Sparse Latent Diffusion
Shang Liu, Chenjie Cao, Chaohui Yu +3
Despite the remarkable developments achieved by recent 3D generation works, scaling these methods to geographic extents, such as modeling thousands of square kilometers of Earth's…
GeoVideo: Introducing Geometric Regularization into Video Generation Model
Yunpeng Bai, Shaoheng Fang, Chaohui Yu +2
Recent advances in video generation have enabled the synthesis of high-quality and visually realistic clips using diffusion transformer models. However, most existing approaches op…
Learning Invariant Representations across Domains and Tasks
Jindong Wang, Wenjie Feng, Chang Liu +5
Being expensive and time-consuming to collect massive COVID-19 image samples to train deep classification models, transfer learning is a promising approach by transferring knowledg…
RynnVLA-002: A Unified Vision-Language-Action and World Model
Jun Cen, Siteng Huang, Yuqian Yuan +11
We introduce RynnVLA-002, a unified Vision-Language-Action (VLA) and world model. The world model leverages action and visual inputs to predict future image states, learning the un…
CamPVG: Camera-Controlled Panoramic Video Generation with Epipolar-Aware Diffusion
Chenhao Ji, Chaohui Yu, Junyao Gao +2
Recently, camera-controlled video generation has seen rapid development, offering more precise control over video generation. However, existing methods predominantly focus on camer…
MVRoom: Controllable 3D Indoor Scene Generation with Multi-View Diffusion Models
Shaoheng Fang, Chaohui Yu, Fan Wang +1
We introduce MVRoom, a controllable novel view synthesis (NVS) pipeline for 3D indoor scenes that uses multi-view diffusion conditioned on a coarse 3D layout. MVRoom employs a two-…
Improved Neural Radiance Fields Using Pseudo-depth and Fusion
Jingliang Li, Qiang Zhou, Chaohui Yu +4
Since the advent of Neural Radiance Fields, novel view synthesis has received tremendous attention. The existing approach for the generalization of radiance field reconstruction pr…
SCOPE: Scale-Consistent One-Pass Estimation of 3D Geometry
Zheng Zhang, Lihe Yang, Tianyu Yang +6
We present SCOPE (Scale-Consistent One-Pass Estimation of 3D Geometry), a novel approach for estimating 3D geometry from extended monocular video sequences, where existing methods…
WorldVLA: Towards Autoregressive Action World Model
Jun Cen, Chaohui Yu, Hangjie Yuan +9
We present WorldVLA, an autoregressive action world model that unifies action and image understanding and generation. Our WorldVLA intergrates Vision-Language-Action (VLA) model an…
VCD-Texture: Variance Alignment based 3D-2D Co-Denoising for Text-Guided Texturing
Shang Liu, Chaohui Yu, Chenjie Cao +2
Recent research on texture synthesis for 3D shapes benefits a lot from dramatically developed 2D text-to-image diffusion models, including inpainting-based and optimization-based a…
Points-to-3D: Bridging the Gap between Sparse Points and Shape-Controllable Text-to-3D Generation
Chaohui Yu, Qiang Zhou, Jingliang Li +3
Text-to-3D generation has recently garnered significant attention, fueled by 2D diffusion models trained on billions of image-text pairs. Existing methods primarily rely on score d…
A Contrastive Pre-trained Foundation Model for Deciphering Imaging Noisomics across Modalities
Yuanjie Gu, Yiqun Wang, Chaohui Yu +4
Characterizing imaging noise is notoriously data-intensive and device-dependent, as modern sensors entangle physical signals with complex algorithmic artifacts. Current paradigms s…
Uni3C: Unifying Precisely 3D-Enhanced Camera and Human Motion Controls for Video Generation
Chenjie Cao, Jingkai Zhou, Shikai Li +5
Camera and human motion controls have been extensively studied for video generation, but existing approaches typically address them separately, suffering from limited data with hig…