papers

Publications (152)

cs.CV2023

Enhancing Quality of Pose-varied Face Restoration with Local Weak Feature Sensing and GAN Prior

Kai Hu, Yu Liu, Renhe Liu +3

Facial semantic guidance (including facial landmarks, facial heatmaps, and facial parsing maps) and facial generative adversarial networks (GAN) prior have been widely used in blin…

cs.CV2022

Hierarchical Normalization for Robust Monocular Depth Estimation

Chi Zhang, Wei Yin, Zhibin Wang +3

In this paper, we address monocular depth estimation with deep neural networks. To enable training of deep monocular estimation models with various sources of datasets, state-of-th…

cs.CV2019

Real-Time Semantic Segmentation via Multiply Spatial Fusion Network

Haiyang Si, Zhiqiang Zhang, Feifan Lv +2

Real-time semantic segmentation plays a significant role in industry applications, such as autonomous driving, robotics and so on. It is a challenging task as both efficiency and p…

eess.AS2026

Step-Audio-R1.5 Technical Report

Yuxin Zhang, Xiangyu Tony Zhang, Daijiao Liu +16

Recent advancements in large audio language models have extended Chain-of-Thought (CoT) reasoning into the auditory domain, enabling models to tackle increasingly complex acoustic…

cs.CV2019

An End-to-End Network for Panoptic Segmentation

Huanyu Liu, Chao Peng, Changqian Yu +4

Panoptic segmentation, which needs to assign a category label to each pixel and segment each object instance simultaneously, is a challenging topic. Traditionally, the existing app…

cs.CL2026

Learning an Efficient Multi-Turn Dialogue Evaluator from Multiple LLM Judges

Yuqi Tang, Kehua Feng, Yunfeng Wang +6

Evaluating the conversational abilities of large language models (LLMs) remains a challenging task. Current mainstream approaches primarily rely on the "LLM-as-a-judge" paradigm, w…

eess.IV2022

Preparing data for pathological artificial intelligence with clinical-grade performance

Yuanqing Yang, Kai Sun, Yanhua Gao +2

[Purpose] The pathology is decisive for disease diagnosis, but relies heavily on the experienced pathologists. Recently, pathological artificial intelligence (PAI) is thought to im…

cs.CV2026

ViStoryBench: Comprehensive Benchmark Suite for Story Visualization

Cailin Zhuang, Ailin Huang, Yaoqi Hu +12

Story visualization aims to generate coherent image sequences that faithfully represent a narrative and match given character references. Despite progress in generative models, exi…

cs.CV2025

Towards Accurate and Interpretable Neuroblastoma Diagnosis via Contrastive Multi-scale Pathological Image Analysis

Zhu Zhu, Shuo Jiang, Jingyuan Zheng +7

Neuroblastoma, adrenal-derived, is among the most common pediatric solid malignancies, characterized by significant clinical heterogeneity. Timely and accurate pathological diagnos…

eess.IV2021

Multi-scale super-resolution generation of low-resolution scanned pathological images

Kai Sun, Yanhua Gao, Ting Xie +5

Background. Digital pathology has aroused widespread interest in modern pathology. The key of digitalization is to scan the whole slide image (WSI) at high magnification. The lager…

cs.CV2023

STAR Loss: Reducing Semantic Ambiguity in Facial Landmark Detection

Zhenglin Zhou, Huaxia Li, Hong Liu +3

Recently, deep learning-based facial landmark detection has achieved significant improvement. However, the semantic ambiguity problem degrades detection performance. Specifically,…

cs.AI2025

Reason from Future: Reverse Thought Chain Enhances LLM Reasoning

Yinlong Xu, Yanzhao Zheng, Shuoshuo Sun +7

It has been demonstrated that carefully designed reasoning paradigms, like Chain-of-Thought (CoT) and Tree-of-Thought (ToT), can enhance the reasoning capabilities of small languag…

cs.CV2017

Light-Head R-CNN: In Defense of Two-Stage Object Detector

Zeming Li, Chao Peng, Gang Yu +3

In this paper, we first investigate why typical two-stage methods are not as fast as single-stage, fast detectors like YOLO and SSD. We find that Faster R-CNN and R-FCN perform an…

cs.CR2010

Provable Secure Identity Based Generalized Signcryption Scheme

Gang Yu, Xiaoxiao Ma, Yong Shen +1

According to actual needs, generalized signcryption scheme can flexibly work as an encryption scheme, a signature scheme or a signcryption scheme. In this paper, firstly, we give a…

cs.CV2021

Object-aware Long-short-range Spatial Alignment for Few-Shot Fine-Grained Image Classification

Yike Wu, Bo Zhang, Gang Yu +4

The goal of few-shot fine-grained image classification is to recognize rarely seen fine-grained objects in the query set, given only a few samples of this class in the support set.…

cs.CV2025

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Guoqing Ma, Haoyang Huang, Kun Yan +112

We present Step-Video-T2V, a state-of-the-art text-to-video pre-trained model with 30B parameters and the ability to generate videos up to 204 frames in length. A deep compression…

cs.CV2018

SFace: An Efficient Network for Face Detection in Large Scale Variations

Jianfeng Wang, Ye Yuan, Boxun Li +2

Face detection serves as a fundamental research topic for many applications like face recognition. Impressive progress has been made especially with the recent development of convo…

cs.CV2025

MotionAgent: Fine-grained Controllable Video Generation via Motion Field Agent

Xinyao Liao, Xianfang Zeng, Liao Wang +3

We propose MotionAgent, enabling fine-grained motion control for text-guided image-to-video generation. The key technique is the motion field agent that converts motion information…

cs.CV2018

MegDet: A Large Mini-Batch Object Detector

Chao Peng, Tete Xiao, Zeming Li +5

The improvements in recent CNN-based object detection works, from R-CNN [11], Fast/Faster R-CNN [10, 31] to recent Mask R-CNN [14] and RetinaNet [24], mainly come from new network,…

cs.CV2019

Class-balanced Grouping and Sampling for Point Cloud 3D Object Detection

Benjin Zhu, Zhengkai Jiang, Xiangxin Zhou +2

This report presents our method which wins the nuScenes3D Detection Challenge [17] held in Workshop on Autonomous Driving(WAD, CVPR 2019). Generally, we utilize sparse 3D convoluti…

cs.CV2025

SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation

Youliang Zhang, Zhaoyang Li, Duomin Wang +6

The rapid development of large-scale models has catalyzed significant breakthroughs in the digital human domain. These advanced methodologies offer high-fidelity solutions for avat…

cs.CV2025

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts

Duomin Wang, Wei Zuo, Aojie Li +7

We introduce UniVerse-1, a unified, Veo-3-like model capable of simultaneously generating coordinated audio and video. To enhance training efficiency, we bypass training from scrat…

cs.CV2026

FreeStyle: Free Control of Style-Content Dual-Reference Generation from Community LoRA Mining

Jinghong Lan, Wei Cheng, Yunuo Chen +10

Style-content dual-reference generation aims to synthesize an image that preserves the structure and semantics of a content reference while adopting the style of a separate style r…

cs.CV2023

VQ-NeRF: Vector Quantization Enhances Implicit Neural Representations

Yiying Yang, Wen Liu, Fukun Yin +4

Recent advancements in implicit neural representations have contributed to high-fidelity surface reconstruction and photorealistic novel view synthesis. However, the computational…

cs.CV2026

GEditBench v2: A Human-Aligned Benchmark for General Image Editing

Zhangqi Jiang, Zheng Sun, Xianfang Zeng +7

Recent advances in image editing have enabled models to handle complex instructions with impressive realism. However, existing evaluation frameworks lag behind: current benchmarks…

cs.CV2025

FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion Understanding

Chongjun Tu, Lin Zhang, Pengtao Chen +5

Multimodal Large Language Models (MLLMs) have shown remarkable capabilities in video content understanding but still struggle with fine-grained motion comprehension. To comprehensi…

eess.SP2024

Synchro-Transient-Extracting Transform for the Analysis of Signals with Both Harmonic and Impulsive Components

Yunlong Ma, Gang Yu, Tianran Lin +1

Time-frequency analysis (TFA) techniques play an important role in the field of machine fault diagnosis attributing to their superiority in dealing with nonstationary signals. Sync…

cs.CV2020

BiSeNet V2: Bilateral Network with Guided Aggregation for Real-time Semantic Segmentation

Changqian Yu, Changxin Gao, Jingbo Wang +3

The low-level details and high-level semantics are both essential to the semantic segmentation task. However, to speed up the model inference, current approaches almost always sacr…

cs.CV2023

Paint3D: Paint Anything 3D with Lighting-Less Texture Diffusion Models

Xianfang Zeng, Xin Chen, Zhongqi Qi +6

This paper presents Paint3D, a novel coarse-to-fine generative framework that is capable of producing high-resolution, lighting-less, and diverse 2K UV texture maps for untextured…

cs.GR2023

TapMo: Shape-aware Motion Generation of Skeleton-free Characters

Jiaxu Zhang, Shaoli Huang, Zhigang Tu +4

Previous motion generation methods are limited to the pre-rigged 3D human model, hindering their applications in the animation of various non-rigged characters. In this work, we pr…

cs.CV2019

Learnable Tree Filter for Structure-preserving Feature Transform

Lin Song, Yanwei Li, Zeming Li +4

Learning discriminative global features plays a vital role in semantic segmentation. And most of the existing methods adopt stacks of local convolutions or non-local blocks to capt…

cs.CV2025

SeaFormer++: Squeeze-enhanced Axial Transformer for Mobile Visual Recognition

Qiang Wan, Zilong Huang, Jiachen Lu +2

Since the introduction of Vision Transformers, the landscape of many computer vision tasks (e.g., semantic segmentation), which has been overwhelmingly dominated by CNNs, recently…

cs.CV2021

Attribute-specific Control Units in StyleGAN for Fine-grained Image Manipulation

Rui Wang, Jian Chen, Gang Yu +4

Image manipulation with StyleGAN has been an increasing concern in recent years.Recent works have achieved tremendous success in analyzing several semantic latent spaces to edit th…

cs.CV2021

Sketch Me A Video

Haichao Zhang, Gang Yu, Tao Chen +1

Video creation has been an attractive yet challenging task for artists to explore. With the advancement of deep learning, recent works try to utilize deep convolutional neural netw…

cs.CV2025

OmniSVG: A Unified Scalable Vector Graphics Generation Model

Yiying Yang, Wei Cheng, Sijin Chen +7

Scalable Vector Graphics (SVG) is an important image format widely adopted in graphic design because of their resolution independence and editability. The study of generating high-…

cs.CV2026

Motion4Motion: Motion Transfer Across Subjects at Inference

Ling-Hao Chen, Zixin Yin, Duomin Wang +2

The paper introduces Motion4Motion, a training‑free framework that transfers motion between videos by modeling motion flow instead of relying on predefined skeletons, enabling tran…

#motion transfer#video animation#cross‑species motion#skeleton‑agnostic
cs.CV2020

Context Prior for Scene Segmentation

Changqian Yu, Jingbo Wang, Changxin Gao +3

Recent works have widely explored the contextual dependencies to achieve more accurate segmentation results. However, most approaches rarely distinguish different types of contextu…

cs.LG2026

SERL: Self-Examining Reinforcement Learning on Open-Domain

Weixuan Ou, Yanzhao Zheng, Shuoshuo Sun +7

Reinforcement Learning (RL) has been shown to improve the capabilities of large language models (LLMs). However, applying RL to open-domain tasks faces two key challenges: (1) the…

cs.CV2026

StyleMe3D: Stylization with Disentangled Priors by Multiple Encoders on 3D Gaussians

Cailin Zhuang, Yaoqi Hu, Xuanyang Zhang +7

Current 3D Gaussian Splatting stylization approaches are limited in their ability to represent diverse artistic styles, frequently defaulting to low-level texture replacement or yi…

cs.CV2023

FaceStudio: Put Your Face Everywhere in Seconds

Yuxuan Yan, Chi Zhang, Rui Wang +5

This study investigates identity-preserving image synthesis, an intriguing task in image generation that seeks to maintain a subject's identity while adding a personalized, stylist…

cs.AI2026

SkillNet: Create, Evaluate, and Connect AI Skills

Yuan Liang, Ruobin Zhong, Haoming Xu +46

Current AI agents can flexibly invoke tools and execute complex tasks, yet their long-term advancement is hindered by the lack of systematic accumulation and transfer of skills. Wi…

cs.CV2019

Rethinking on Multi-Stage Networks for Human Pose Estimation

Wenbo Li, Zhicheng Wang, Binyi Yin +7

Existing pose estimation approaches fall into two categories: single-stage and multi-stage methods. While multi-stage methods are seemingly more suited for the task, their performa…

cs.CV2025

In-Context Learning with Unpaired Clips for Instruction-based Video Editing

Xinyao Liao, Xianfang Zeng, Ziye Song +3

Despite the rapid progress of instruction-based image editing, its extension to video remains underexplored, primarily due to the prohibitive cost and complexity of constructing la…

cs.CV2024

Generative Motion Stylization of Cross-structure Characters within Canonical Motion Space

Jiaxu Zhang, Xin Chen, Gang Yu +1

Stylized motion breathes life into characters. However, the fixed skeleton structure and style representation hinder existing data-driven motion synthesis methods from generating s…

cs.CV2023

Metric3D: Towards Zero-shot Metric 3D Prediction from A Single Image

Wei Yin, Chi Zhang, Hao Chen +5

Reconstructing accurate 3D scenes from images is a long-standing vision task. Due to the ill-posedness of the single-image reconstruction problem, most well-established methods are…

cs.CV2023

Robust Geometry-Preserving Depth Estimation Using Differentiable Rendering

Chi Zhang, Wei Yin, Gang Yu +5

In this study, we address the challenge of 3D scene structure recovery from monocular depth estimation. While traditional depth estimation methods leverage labeled datasets to dire…

cs.CL2026

Chronological Thinking in Full-Duplex Spoken Dialogue Language Models

Donghang Wu, Haoyang Zhang, Chen Chen +8

Recent advances in spoken dialogue language models (SDLMs) reflect growing interest in shifting from turn-based to full-duplex systems, where the models continuously perceive user…

cs.CV2025

ReasonEdit: Towards Reasoning-Enhanced Image Editing Models

Fukun Yin, Shiyu Liu, Yucheng Han +12

Recent advances in image editing models have shown remarkable progress. A common architectural design couples a multimodal large language model (MLLM) encoder with a diffusion deco…

cs.CV2023

Disentangled Pre-training for Image Matting

Yanda Li, Zilong Huang, Gang Yu +3

Image matting requires high-quality pixel-level human annotations to support the training of a deep model in recent literature. Whereas such annotation is costly and hard to scale,…

cs.CL2026

Rethinking Memory as Continuously Evolving Connectivity

Jizhan Fang, Buqiang Xu, Zhixian Wang +12

Existing memory-augmented LLM agents often treat memory as a static repository with pre-defined representations and fixed retrieval pipelines, which is brittle in dynamic agentic e…

cs.CV2023

Capturing the motion of every joint: 3D human pose and shape estimation with independent tokens

Sen Yang, Wen Heng, Gang Liu +3

In this paper we present a novel method to estimate 3D human pose and shape from monocular videos. This task requires directly recovering pixel-alignment 3D human pose and body sha…

cs.CV2024

Lightweight Model Pre-training via Language Guided Knowledge Distillation

Mingsheng Li, Lin Zhang, Mingzhen Zhu +4

This paper studies the problem of pre-training for small models, which is essential for many mobile devices. Current state-of-the-art methods on this problem transfer the represent…

cs.CV2020

SiamFC++: Towards Robust and Accurate Visual Tracking with Target Estimation Guidelines

Yinda Xu, Zeyu Wang, Zuoxin Li +2

Visual tracking problem demands to efficiently perform robust classification and accurate target state estimation over a given target at the same time. Former methods have proposed…

cs.AI2025

Step-Audio-R1 Technical Report

Fei Tian, Xiangyu Tony Zhang, Yuxin Zhang +14

Recent advances in reasoning models have demonstrated remarkable success in text and vision domains through extended chain-of-thought deliberation. However, a perplexing phenomenon…

cs.CV2023

MotionGPT: Human Motion as a Foreign Language

Biao Jiang, Xin Chen, Wen Liu +3

Though the advancement of pre-trained large language models unfolds, the exploration of building a unified model for language and other multi-modal data, such as motion, remains ch…

cs.CV2020

State-Aware Tracker for Real-Time Video Object Segmentation

Xi Chen, Zuoxin Li, Ye Yuan +3

In this work, we address the task of semi-supervised video object segmentation(VOS) and explore how to make efficient use of video property to tackle the challenge of semi-supervis…

cs.CV2022

Efficient Single-Image Depth Estimation on Mobile Devices, Mobile AI & AIM 2022 Challenge: Report

Andrey Ignatov, Grigory Malivenko, Radu Timofte +36

Various depth estimation models are now widely used on many mobile and IoT devices for image segmentation, bokeh effect rendering, object tracking and many other mobile tasks. Thus…

eess.AS2026

StepAudio 2.5 Technical Report

Bin Lin, Bo Zhao, Boyong Wu +98

Unified audio-language modeling has emerged as a prominent trend in modern speech systems, promising to bring the reasoning capabilities of large language models to auditory tasks.…

math.NT2006

Sieving by large integers and covering systems of congruences

Michael Filaseta, Kevin Ford, Sergei Konyagin +2

An old question of Erdos asks if there exists, for each number N, a finite set S of integers greater than N and residue classes r(n) mod n for n in S whose union is all the integer…

cs.CV2023

StableLLaVA: Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data

Yanda Li, Chi Zhang, Gang Yu +6

The remarkable multimodal capabilities demonstrated by OpenAI's GPT-4 have sparked significant interest in the development of multimodal Large Language Models (LLMs). A primary res…

eess.SP2026

High-order synchrosqueezed wavelet-chirplet transform for instantaneous frequency and chirprate estimation

Shuixin Li, Jiecheng Chen, Qingtang Jiang +1

The separation of multicomponent signals with crossing instantaneous frequency (IF) curves remains a fundamental challenge in time-frequency analysis. Although the synchrosqueezed…

cs.CV2018

Modeling Local Geometric Structure of 3D Point Clouds using Geo-CNN

Shiyi Lan, Ruichi Yu, Gang Yu +1

Recent advances in deep convolutional neural networks (CNNs) have motivated researchers to adapt CNNs to directly model points in 3D point clouds. Modeling local structure has been…

cs.CV2025

DreamDance: Animating Character Art via Inpainting Stable Gaussian Worlds

Jiaxu Zhang, Xianfang Zeng, Xin Chen +4

This paper presents DreamDance, a novel character art animation framework capable of producing stable, consistent character and scene motion conditioned on precise camera trajector…

cs.CV2024

Cross-Dimensional Medical Self-Supervised Representation Learning Based on a Pseudo-3D Transformation

Fei Gao, Siwen Wang, Fandong Zhang +5

Medical image analysis suffers from a shortage of data, whether annotated or not. This becomes even more pronounced when it comes to 3D medical images. Self-Supervised Learning (SS…

cs.CV2026

AppAgent: Multimodal Agents as Smartphone Users

Chi Zhang, Zhao Yang, Jiaxuan Liu +6

Recent advancements in large language models (LLMs) have led to the creation of intelligent agents capable of performing complex tasks. This paper introduces a novel LLM-based mult…

cs.CV2018

Scene Text Detection with Supervised Pyramid Context Network

Enze Xie, Yuhang Zang, Shuai Shao +3

Scene text detection methods based on deep learning have achieved remarkable results over the past years. However, due to the high diversity and complexity of natural scenes, previ…

cs.CV2025

Step1X-3D: Towards High-Fidelity and Controllable Generation of Textured 3D Assets

Weiyu Li, Xuanyang Zhang, Zheng Sun +15

While generative artificial intelligence has advanced significantly across text, image, audio, and video domains, 3D generation remains comparatively underdeveloped due to fundamen…

cs.CV2024

MeshXL: Neural Coordinate Field for Generative 3D Foundation Models

Sijin Chen, Xin Chen, Anqi Pang +11

The polygon mesh representation of 3D data exhibits great flexibility, fast rendering speed, and storage efficiency, which is widely preferred in various applications. However, giv…

cs.CV2020

High-Order Information Matters: Learning Relation and Topology for Occluded Person Re-Identification

Guan'an Wang, Shuo Yang, Huanyu Liu +6

Occluded person re-identification (ReID) aims to match occluded person images to holistic ones across dis-joint cameras. In this paper, we propose a novel framework by learning hig…

cs.CV2026

PixelSmile: Toward Fine-Grained Facial Expression Editing

Jiabin Hua, Hengyuan Xu, Aojie Li +4

Fine-grained facial expression editing has long been limited by intrinsic semantic overlap. To address this, we construct the Flex Facial Expression (FFE) dataset with continuous a…

cs.CV2019

TACNet: Transition-Aware Context Network for Spatio-Temporal Action Detection

Lin Song, Shiwei Zhang, Gang Yu +1

Current state-of-the-art approaches for spatio-temporal action detection have achieved impressive results but remain unsatisfactory for temporal extent detection. The main reason c…

cs.CV2018

Cascaded Pyramid Network for Multi-Person Pose Estimation

Yilun Chen, Zhicheng Wang, Yuxiang Peng +3

The topic of multi-person pose estimation has been largely improved recently, especially with the development of convolutional neural network. However, there still exist a lot of c…

cs.CV2025

Native 3D Editing with Full Attention

Weiwei Cai, Shuangkang Fang, Weicai Ye +7

Instruction-guided 3D editing is a rapidly emerging field with the potential to broaden access to 3D content creation. However, existing methods face critical limitations: optimiza…

cs.CV2023

IT3D: Improved Text-to-3D Generation with Explicit View Synthesis

Yiwen Chen, Chi Zhang, Xiaofeng Yang +4

Recent strides in Text-to-3D techniques have been propelled by distilling knowledge from powerful large text-to-image diffusion models (LDMs). Nonetheless, existing Text-to-3D appr…

cs.CV2018

DetNet: A Backbone network for Object Detection

Zeming Li, Chao Peng, Gang Yu +3

Recent CNN based object detectors, no matter one-stage methods like YOLO, SSD, and RetinaNe or two-stage detectors like Faster R-CNN, R-FCN and FPN are usually trying to directly f…

cs.CV2025

Metric3Dv2: A Versatile Monocular Geometric Foundation Model for Zero-shot Metric Depth and Surface Normal Estimation

Mu Hu, Wei Yin, Chi Zhang +7

We introduce Metric3D v2, a geometric foundation model for zero-shot metric depth and surface normal estimation from a single image, which is crucial for metric 3D recovery. While…

cs.CV2023

ChartLlama: A Multimodal LLM for Chart Understanding and Generation

Yucheng Han, Chi Zhang, Xin Chen +5

Multi-modal large language models have demonstrated impressive performances on most vision-language tasks. However, the model generally lacks the understanding capabilities for spe…

cs.CV2026

Vision Foundation Models as Generalist Tokenizers for Image Generation

Anlin Zheng, Qi Han, Xin Wen +5

In this work, we explore the largely unexplored direction of building a generalist image tokenizer directly on top of a frozen vision foundation model (VFM). To build this tokenize…

cs.CV2023

ShapeGPT: 3D Shape Generation with A Unified Multi-modal Language Model

Fukun Yin, Xin Chen, Chi Zhang +6

The advent of large language models, enabling flexibility through instruction-driven approaches, has revolutionized many traditional generative tasks, but large models for 3D data,…

cs.CV2019

WIDER Face and Pedestrian Challenge 2018: Methods and Results

Chen Change Loy, Dahua Lin, Wanli Ouyang +49

This paper presents a review of the 2018 WIDER Challenge on Face and Pedestrian. The challenge focuses on the problem of precise localization of human faces and bodies, and accurat…

cs.CL2026

Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks

Tianze Xu, Yanzhao Zheng, Pengrui Lu +11

Rubric-based Reinforcement Learning (RL) has emerged as a promising approach for aligning Large Language Models (LLMs) with complex, open-domain instruction following tasks. Howeve…

cs.CV2019

Double Anchor R-CNN for Human Detection in a Crowd

Kevin Zhang, Feng Xiong, Peize Sun +3

Detecting human in a crowd is a challenging problem due to the uncertainties of occlusion patterns. In this paper, we propose to handle the crowd occlusion problem in human detecti…

cs.CV2017

Face Attention Network: An Effective Face Detector for the Occluded Faces

Jianfeng Wang, Ye Yuan, Gang Yu

The performance of face detection has been largely improved with the development of convolutional neural network. However, the occlusion issue due to mask and sunglasses, is still…

cs.CV2024

MeshAnything: Artist-Created Mesh Generation with Autoregressive Transformers

Yiwen Chen, Tong He, Di Huang +9

Recently, 3D assets created via reconstruction and generation have matched the quality of manually crafted assets, highlighting their potential for replacement. However, this poten…

astro-ph.HE2020

Reconstruction of Supernova Gravitational Waves Waveforms: Comparing Three Time-frequency Transform Methods

Zhuotao Li, Xilong Fan, Gang Yu

For supernovae gravitational wave signal analysis which intend to reconstruct supernova gravitational waves waveforms, we compare the performance of short-time Fourier transform (S…

cs.LG2022

Designing thermal radiation metamaterials via hybrid adversarial autoencoder and Bayesian optimization

Dezhao Zhu, Jiang Guo, Gang Yu +3

Designing thermal radiation metamaterials is challenging especially for problems with high degrees of freedom and complex objective. In this letter, we have developed a hybrid mate…

cs.CV2019

Shape Robust Text Detection with Progressive Scale Expansion Network

Wenhai Wang, Enze Xie, Xiang Li +4

Scene text detection has witnessed rapid progress especially with the recent development of convolutional neural networks. However, there still exists two challenges which prevent…

cs.CL2025

Step-Audio 2 Technical Report

Boyong Wu, Chao Yan, Chen Hu +106

This paper presents Step-Audio 2, an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation. By integrating a latent…

cs.LG2026

SkillRouter: Skill Routing for LLM Agents at Scale

YanZhao Zheng, ZhenTao Zhang, Chao Ma +8

Reusable skills let LLM agents package task-specific procedures, tool affordances, and execution guidance into modular building blocks. As skill ecosystems grow to tens of thousand…

cs.LG2025

DeRS: Towards Extremely Efficient Upcycled Mixture-of-Experts Models

Yongqi Huang, Peng Ye, Chenyu Huang +5

Upcycled Mixture-of-Experts (MoE) models have shown great potential in various tasks by converting the original Feed-Forward Network (FFN) layers in pre-trained dense models into M…

cs.CV2024

MikuDance: Animating Character Art with Mixed Motion Dynamics

Jiaxu Zhang, Xianfang Zeng, Xin Chen +3

We propose MikuDance, a diffusion-based pipeline incorporating mixed motion dynamics to animate stylized character art. MikuDance consists of two key techniques: Mixed Motion Model…

cs.CL2025

Step-Audio-EditX Technical Report

Chao Yan, Boyong Wu, Peng Yang +12

We present Step-Audio-EditX, the first open-source LLM-based audio model excelling at expressive and iterative audio editing encompassing emotion, speaking style, and paralinguisti…

cs.CV2020

Efficient and Accurate Arbitrary-Shaped Text Detection with Pixel Aggregation Network

Wenhai Wang, Enze Xie, Xiaoge Song +5

Scene text detection, an important step of scene text reading systems, has witnessed rapid development with convolutional neural networks. Nonetheless, two main challenges still ex…

cs.CV2024

Scene123: One Prompt to 3D Scene Generation via Video-Assisted and Consistency-Enhanced MAE

Yiying Yang, Fukun Yin, Jiayuan Fan +3

As Artificial Intelligence Generated Content (AIGC) advances, a variety of methods have been developed to generate text, images, videos, and 3D objects from single or multimodal in…

cs.CV2023

M3DBench: Let's Instruct Large Models with Multi-modal 3D Prompts

Mingsheng Li, Xin Chen, Chi Zhang +5

Recently, 3D understanding has become popular to facilitate autonomous agents to perform further decisionmaking. However, existing 3D datasets and methods are often limited to spec…

cs.CV2023

Michelangelo: Conditional 3D Shape Generation based on Shape-Image-Text Aligned Latent Representation

Zibo Zhao, Wen Liu, Xin Chen +7

We present a novel alignment-before-generation approach to tackle the challenging task of generating general 3D shapes based on 2D images or texts. Directly learning a conditional…

cs.CV2018

BiSeNet: Bilateral Segmentation Network for Real-time Semantic Segmentation

Changqian Yu, Jingbo Wang, Chao Peng +3

Semantic segmentation requires both rich spatial information and sizeable receptive field. However, modern approaches usually compromise spatial resolution to achieve real-time inf…

cs.CV2022

Learning Variational Motion Prior for Video-based Motion Capture

Xin Chen, Zhuo Su, Lingbo Yang +4

Motion capture from a monocular video is fundamental and crucial for us humans to naturally experience and interact with each other in Virtual Reality (VR) and Augmented Reality (A…

cs.CV2026

Head Forcing: Long Autoregressive Video Generation via Head Heterogeneity

Jiahao Tian, Yiwei Wang, Gang Yu +1

Autoregressive video diffusion models support real-time synthesis but suffer from error accumulation and context loss over long horizons. We discover that attention heads in AR vid…

cs.CL2026

Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models

Donghang Wu, Haoyang Zhang, Jun Chen +9

Real-time Spoken Language Models (SLMs) struggle to leverage Chain-of-Thought (CoT) reasoning due to the prohibitive latency of generating the entire thought process sequentially.…