papers

Publications (259)

cs.CV2020

Word-level Deep Sign Language Recognition from Video: A New Large-scale Dataset and Methods Comparison

Dongxu Li, Cristian Rodriguez Opazo, Xin Yu +1

Vision-based sign language recognition aims at helping deaf people to communicate with others. However, most existing sign language datasets are limited to a small number of words.…

cs.LG2025

Virtual Width Networks

Seed, Baisheng Li, Banggu Wu +115

We introduce Virtual Width Networks (VWN), a framework that delivers the benefits of wider representations without incurring the quadratic cost of increasing the hidden size. VWN d…

cs.CL2026

TAMTRL: Teacher-Aligned Reward Reshaping for Multi-Turn Reinforcement Learning in Long-Context Compression

Li Wang, Yandong Wang, Xin Yu +3

The rapid progress of large language models (LLMs) has led to remarkable performance gains across a wide range of tasks. However, when handling long documents that exceed the model…

cs.CV2024

PlantSeg: A Large-Scale In-the-wild Dataset for Plant Disease Segmentation

Tianqi Wei, Zhi Chen, Xin Yu +3

Plant diseases pose significant threats to agriculture. It necessitates proper diagnosis and effective treatment to safeguard crop yields. To automate the diagnosis process, image…

cs.CV2021

Gait Recognition via Effective Global-Local Feature Representation and Local Temporal Aggregation

Beibei Lin, Shunli Zhang, Xin Yu

Gait recognition is one of the most important biometric technologies and has been applied in many fields. Recent gait recognition frameworks represent each gait frame by descriptor…

hep-ph2014

Semileptonic decays in the perturbative QCD approach

Wen-Fei Wang, Xin Yu, Cai-Dian Lü +1

In this paper we study the semileptonic decays of (here stands for , , or ). After evaluating the $B_c^+ \to (D_{(s…

eess.IV2023

Multi-Contrast Computed Tomography Atlas of Healthy Pancreas

Yinchi Zhou, Ho Hin Lee, Yucheng Tang +6

With the substantial diversity in population demographics, such as differences in age and body composition, the volumetric morphology of pancreas varies greatly, resulting in disti…

cs.CV2021

Joint 3D Human Shape Recovery and Pose Estimation from a Single Image with Bilayer Graph

Xin Yu, Jeroen van Baar, Siheng Chen

The ability to estimate the 3D human shape and pose from images can be useful in many contexts. Recent approaches have explored using graph convolutional networks and achieved prom…

cs.AI2026

Unsupervised Adaptation of PDE Foundation Models

Ziye Song, Zhao Wei, Xin Yu +2

Pretrained partial differential equation (PDE) foundation models can generalize across different equations, but adapting them to unseen PDE systems typically requires dense solutio…

cs.CV2025

Trust-Aware Diversion for Data-Effective Distillation

Zhuojie Wu, Yanbin Liu, Xin Shen +2

Dataset distillation compresses a large dataset into a small synthetic subset that retains essential information. Existing methods assume that all samples are perfectly labeled, li…

cs.CV2020

Copy and Paste GAN: Face Hallucination from Shaded Thumbnails

Yang Zhang, Ivor Tsang, Yawei Luo +3

Existing face hallucination methods based on convolutional neural networks (CNN) have achieved impressive performance on low-resolution (LR) faces in a normal illumination conditio…

cs.LG2024

Machine Unlearning via Null Space Calibration

Huiqiang Chen, Tianqing Zhu, Xin Yu +1

Machine unlearning aims to enable models to forget specific data instances when receiving deletion requests. Current research centres on efficient unlearning to erase the influence…

cs.CV2021

ARVo: Learning All-Range Volumetric Correspondence for Video Deblurring

Dongxu Li, Chenchen Xu, Kaihao Zhang +5

Video deblurring models exploit consecutive frames to remove blurs from camera shakes and object motions. In order to utilize neighboring sharp patches, typical methods rely mainly…

math.AP2025

Observation estimates for a semilinear heat equation in \mathbb{R}^n

Guojie Zheng, Xin Yu

This paper studies the state observation problems for the semilinear heat equation in R^n. We derive observation estimates for the equation using the logarithmic convexity property…

stat.ML2025

Understanding the Statistical Accuracy-Communication Trade-off in Personalized Federated Learning with Minimax Guarantees

Xin Yu, Zelin He, Ying Sun +2

Personalized federated learning (PFL) offers a flexible framework for aggregating information across distributed clients with heterogeneous data. This work considers a personalized…

cs.CV2022

Deep Idempotent Network for Efficient Single Image Blind Deblurring

Yuxin Mao, Zhexiong Wan, Yuchao Dai +1

Single image blind deblurring is highly ill-posed as neither the latent sharp image nor the blur kernel is known. Even though considerable progress has been made, several major dif…

cs.LG2026

EXACT: Explicit Attribute-Guided Decoding-Time Personalization

Xin Yu, Hanwen Xing, Lingzhou Xue

Achieving personalized alignment requires adapting large language models to each user's evolving context. While decoding-time personalization offers a scalable alternative to train…

cs.CV2021

VidFace: A Full-Transformer Solver for Video FaceHallucination with Unaligned Tiny Snapshots

Yuan Gan, Yawei Luo, Xin Yu +2

In this paper, we investigate the task of hallucinating an authentic high-resolution (HR) human face from multiple low-resolution (LR) video snapshots. We propose a pure transforme…

cs.CV2026

LightMover: Generative Light Movement with Color and Intensity Controls

Gengze Zhou, Tianyu Wang, Soo Ye Kim +7

We present LightMover, a framework for controllable light manipulation in single images that leverages video diffusion priors to produce physically plausible illumination changes w…

cs.CV2023

When 3D Bounding-Box Meets SAM: Point Cloud Instance Segmentation with Weak-and-Noisy Supervision

Qingtao Yu, Heming Du, Chen Liu +1

Learning from bounding-boxes annotations has shown great potential in weakly-supervised 3D point cloud instance segmentation. However, we observed that existing methods would suffe…

math.OC2018

Estimating the Distribution of Random Parameters in a Diffusion Equation Forward Model for a Transdermal Alcohol Biosensor

Melike Sirlanci, Susan E. Luczak, Catharine E. Fairbairn +4

We estimate the distribution of random parameters in a distributed parameter model with unbounded input and output for the transdermal transport of ethanol in humans. The model tak…

eess.IV2024

Enhancing Single-Slice Segmentation with 3D-to-2D Unpaired Scan Distillation

Xin Yu, Qi Yang, Han Liu +10

2D single-slice abdominal computed tomography (CT) enables the assessment of body habitus and organ health with low radiation exposure. However, single-slice data necessitates the…

cs.MA2026

Modeling Earth-Scale Human-Like Societies with One Billion Agents

Haoxiang Guan, Jiyan He, Liyang Fan +10

Understanding the dynamic evolution of complex social phenomena requires both high-fidelity modeling of human behavior and large-scale simulations. Traditional agent-based models (…

cs.CV2024

Benchmarking In-the-wild Multimodal Disease Recognition and A Versatile Baseline

Tianqi Wei, Zhi Chen, Zi Huang +1

Existing plant disease classification models have achieved remarkable performance in recognizing in-laboratory diseased images. However, their performance often significantly degra…

cs.CV2023

OpenSight: A Simple Open-Vocabulary Framework for LiDAR-Based Object Detection

Hu Zhang, Jianhua Xu, Tao Tang +4

Traditional LiDAR-based object detection research primarily focuses on closed-set scenarios, which falls short in complex real-world applications. Directly transferring existing 2D…

cs.CE2022

Surrogate Neural Network Model for Sensitivity Analysis and Uncertainty Quantification of the Mechanical Behavior in the Optical Lens-Barrel Assembly

Shantanu Shahane, Erman Guleryuz, Diab W Abueidda +7

Surrogate neural network-based models have been lately trained and used in a variety of science and engineering applications where the number of evaluations of a target function is…

math.AP2014

Global existence of null-form wave equations on small asymptotically Euclidean manifolds

Chengbo Wang, Xin Yu

We prove the global existence of the small solutions to the Cauchy problem for quasilinear wave equations satisfying the null condition on , where the metric is a sma…

cond-mat.str-el2026

Observation of Kondo hybridization wave in UTe2

Xin Yu, Shuikang Yu, Zheyu Wu +9

Condensed matter systems with strong electronic correlations often manifest a variety of intertwined ordered phases of charge, spin, orbital and other degrees of freedom. As a prot…

math.PR2018

Generalized Lyapunov criteria on finite-time stability of stochastic nonlinear systems

Xin Yu, Juliang Yin, Suiyang Khoo

This paper considers the problem of finite-time stability for stochastic nonlinear systems. A new Lyapunov theorem of stochastic finite-time stability is proposed, and an important…

cs.CV2023

Exploring Active 3D Object Detection from a Generalization Perspective

Yadan Luo, Zhuoxiao Chen, Zijian Wang +3

To alleviate the high annotation cost in LiDAR-based 3D object detection, active learning is a promising solution that learns to select only a small portion of unlabeled data to an…

cs.CV2023

Hybrid Neural Rendering for Large-Scale Scenes with Motion Blur

Peng Dai, Yinda Zhang, Xin Yu +2

Rendering novel view images is highly desirable for many applications. Despite recent progress, it remains challenging to render high-fidelity and view-consistent novel views of la…

eess.IV2023

Scaling Up 3D Kernels with Bayesian Frequency Re-parameterization for Medical Image Segmentation

Ho Hin Lee, Quan Liu, Shunxing Bao +7

With the inspiration of vision transformers, the concept of depth-wise convolution revisits to provide a large Effective Receptive Field (ERF) using Large Kernel (LK) sizes for med…

cs.CV2019

Learning Strict Identity Mappings in Deep Residual Networks

Xin Yu, Zhiding Yu, Srikumar Ramalingam

A family of super deep networks, referred to as residual networks or ResNet, achieved record-beating performance in various visual tasks such as image recognition, object detection…

cs.CV2024

Affective Behaviour Analysis via Integrating Multi-Modal Knowledge

Wei Zhang, Feng Qiu, Chen Liu +4

Affective Behavior Analysis aims to facilitate technology emotionally smart, creating a world where devices can understand and react to our emotions as humans do. To comprehensivel…

cs.CV2023

RVD: A Handheld Device-Based Fundus Video Dataset for Retinal Vessel Segmentation

MD Wahiduzzaman Khan, Hongwei Sheng, Hu Zhang +11

Retinal vessel segmentation is generally grounded in image-based datasets collected with bench-top devices. The static images naturally lose the dynamic characteristics of retina f…

cs.CV2024

CF-PRNet: Coarse-to-Fine Prototype Refining Network for Point Cloud Completion and Reconstruction

Zhi Chen, Tianqi Wei, Zecheng Zhao +6

In modern agriculture, precise monitoring of plants and fruits is crucial for tasks such as high-throughput phenotyping and automated harvesting. This paper addresses the challenge…

cs.CV2023

Autonomous Stabilization of Retinal Videos for Streamlining Assessment of Spontaneous Venous Pulsations

Hongwei Sheng, Xin Yu, Feiyu Wang +4

Spontaneous retinal Venous Pulsations (SVP) are rhythmic changes in the caliber of the central retinal vein and are observed in the optic disc region (ODR) of the retina. Its absen…

cs.CV2020

Learning Object Relation Graph and Tentative Policy for Visual Navigation

Heming Du, Xin Yu, Liang Zheng

Target-driven visual navigation aims at navigating an agent towards a given target based on the observation of the agent. In this task, it is critical to learn informative visual r…

cs.CL2026

CurveShift: Is Agent Progress Scalar? Separating Level from Shape

Hanwen Xing, Pengyun Wang, BingXu Meng +8

Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate benchmark score. These summa…

cs.LG2026

Explainable LLM Unlearning Through Reasoning

Junfeng Liao, Qizhou Wang, Shanshan Ye +3

LLM unlearning is essential for mitigating safety, copyright, and privacy concerns in pre-trained large language models (LLMs). Compared to preference alignment, it offers a more e…

cs.CV2024

EfficientDreamer: High-Fidelity and Robust 3D Creation via Orthogonal-view Diffusion Prior

Zhipeng Hu, Minda Zhao, Chaoyi Zhao +6

While image diffusion models have made significant progress in text-driven 3D content creation, they often fail to accurately capture the intended meaning of text prompts, especial…

cs.CV2020

TSPNet: Hierarchical Feature Learning via Temporal Semantic Pyramid for Sign Language Translation

Dongxu Li, Chenchen Xu, Xin Yu +4

Sign language translation (SLT) aims to interpret sign video sequences into text-based natural language sentences. Sign videos consist of continuous sequences of sign gestures with…

cs.CV2023

Texture Generation on 3D Meshes with Point-UV Diffusion

Xin Yu, Peng Dai, Wenbo Li +3

In this work, we focus on synthesizing high-quality textures on 3D meshes. We present Point-UV diffusion, a coarse-to-fine pipeline that marries the denoising diffusion model with…

eess.IV2022

Reducing Positional Variance in Cross-sectional Abdominal CT Slices with Deep Conditional Generative Models

Xin Yu, Qi Yang, Yucheng Tang +8

2D low-dose single-slice abdominal computed tomography (CT) slice enables direct measurements of body composition, which are critical to quantitatively characterizing health relati…

cs.CV2019

Optimal Feature Transport for Cross-View Image Geo-Localization

Yujiao Shi, Xin Yu, Liu Liu +2

This paper addresses the problem of cross-view image geo-localization, where the geographic location of a ground-level street-view query image is estimated by matching it against a…

cs.CL2026

Beyond Static Dialogues: Benchmarking Realistic, Heterogeneous, and Evolving Long-Term Memory

Han Zhang, Zihao Tang, Xin Yu +8

In existing memory benchmarks for Large Language Models (LLMs), the evaluated dialogue sessions often lack long-term semantic consistency, and the underlying personas tend to be fl…

eess.IV2025

CMamba: Learned Image Compression with State Space Models

Zhuojie Wu, Heming Du, Shuyun Wang +4

Learned Image Compression (LIC) has explored various architectures, such as Convolutional Neural Networks (CNNs) and transformers, in modeling image content distributions in order…

nlin.SI2018

Periodic parabola solitons for the nonautonomous KP equation

Yingyou Ma, Zhiqiang Chen, Xin Yu

Kadomtsev-Petviashvili (KP) equation, who can describe different models in fluids and plasmas, has drawn investigation for its solitonic solutions with various methods. In this pap…

math.AP2011

Generalized Strichartz Estimates on Perturbed Wave Equation and Applications on Strauss Conjecture

Xin Yu

In this paper we show a general Strichartz estimate for certain perturbed wave equation, and here we can drop the nontrapping hypothesis and handle trapping obstacles with some los…

cs.LG2025

Analytical Survey of Learning with Low-Resource Data: From Analysis to Investigation

Xiaofeng Cao, Mingwei Xu, Xin Yu +8

Learning with high-resource data has demonstrated substantial success in artificial intelligence (AI); however, the costs associated with data annotation and model training remain…

eess.IV2022

Single Slice Thigh CT Muscle Group Segmentation with Domain Adaptation and Self-Training

Qi Yang, Xin Yu, Ho Hin Lee +8

Objective: Thigh muscle group segmentation is important for assessment of muscle anatomy, metabolic disease and aging. Many efforts have been put into quantifying muscle tissues wi…

cs.CV2022

Learning Implicit Body Representations from Double Diffusion Based Neural Radiance Fields

Guangming Yao, Hongzhi Wu, Yi Yuan +3

In this paper, we present a novel double diffusion based neural radiance field, dubbed DD-NeRF, to reconstruct human body geometry and render the human body appearance in novel vie…

cs.CL2024

RichRAG: Crafting Rich Responses for Multi-faceted Queries in Retrieval-Augmented Generation

Shuting Wang, Xin Yu, Mang Wang +3

Retrieval-augmented generation (RAG) effectively addresses issues of static knowledge and hallucination in large language models. Existing studies mostly focus on question scenario…

cs.LG2021

Scaling Up Exact Neural Network Compression by ReLU Stability

Thiago Serra, Xin Yu, Abhinav Kumar +1

We can compress a rectifier network while exactly preserving its underlying functionality with respect to a given input domain if some of its neurons are stable. However, current a…

cs.CV2025

UniTok: A Unified Tokenizer for Visual Generation and Understanding

Chuofan Ma, Yi Jiang, Junfeng Wu +5

Visual generative and understanding models typically rely on distinct tokenizers to process images, presenting a key challenge for unifying them within a single framework. Recent s…

eess.IV2024

Enhancing Hierarchical Transformers for Whole Brain Segmentation with Intracranial Measurements Integration

Xin Yu, Yucheng Tang, Qi Yang +4

Whole brain segmentation with magnetic resonance imaging (MRI) enables the non-invasive measurement of brain regions, including total intracranial volume (TICV) and posterior fossa…

physics.soc-ph2026

Effective Graph Resistance as Cumulative Heat Dissipation

Xiangrong Wang, Xin Yu, Zongze Wu +1

Effective graph resistance is a fundamental structural metric in network science, widely used to quantify global connectivity, compare network architectures, and assess robustness…

cs.LG2025

NVIDIA Nemotron Parse 1.1

Kateryna Chumachenko, Amala Sanjay Deshmukh, Jarno Seppanen +30

We introduce Nemotron-Parse-1.1, a lightweight document parsing and OCR model that advances the capabilities of its predecessor, Nemoretriever-Parse-1.0. Nemotron-Parse-1.1 deliver…

cs.CV2022

Instance As Identity: A Generic Online Paradigm for Video Instance Segmentation

Feng Zhu, Zongxin Yang, Xin Yu +2

Modeling temporal information for both detection and tracking in a unified framework has been proved a promising solution to video instance segmentation (VIS). However, how to effe…

cs.LG2022

The Combinatorial Brain Surgeon: Pruning Weights That Cancel One Another in Neural Networks

Xin Yu, Thiago Serra, Srikumar Ramalingam +1

Neural networks tend to achieve better accuracy with training if they are larger -- even if the resulting models are overparameterized. Nevertheless, carefully removing such excess…

math.AP2011

Concerning the Strauss conjecture on asymptotically Euclidean manifolds

Chengbo Wang, Xin Yu

In this paper we verify the Strauss conjecture for semilinear wave equations on asymptotically Euclidean manifolds when n=3,4, we also give an almost sharp life span for the subcri…

cs.CV2020

Where am I looking at? Joint Location and Orientation Estimation by Cross-View Matching

Yujiao Shi, Xin Yu, Dylan Campbell +1

Cross-view geo-localization is the problem of estimating the position and orientation (latitude, longitude and azimuth angle) of a camera at ground level given a large-scale databa…

hep-ph2012

Charmed Scalar Meson Production in Decays

Yue-Long Shen, Xin Yu

The study on the charmed scalar meson spectroscopy has become a hot topic both experimentally and theoretically. The decays provide an ideal place to study their property…

math.AP2011

Generalized and weighted Strichartz estimates

Jin-Cheng Jiang, Chengbo Wang, Xin Yu

In this paper, we explore the relations between different kinds of Strichartz estimates and give new estimates in Euclidean space . In particular, we prove the genera…

cs.LG2025

The Sound of Risk: A Multimodal Physics-Informed Acoustic Model for Forecasting Market Volatility and Enhancing Market Interpretability

Xiaoliang Chen, Xin Yu, Le Chang +8

Information asymmetry in financial markets, often amplified by strategically crafted corporate narratives, undermines the effectiveness of conventional textual analysis. We propose…

cs.MA2025

Heterogeneity in Multi-Agent Reinforcement Learning

Tianyi Hu, Zhiqiang Pu, Yuan Wang +3

Heterogeneity is a fundamental property in multi-agent reinforcement learning (MARL), which is closely related not only to the functional differences of agents, but also to policy…

cs.CV2020

Weakly-Supervised Salient Object Detection via Scribble Annotations

Jing Zhang, Xin Yu, Aixuan Li +3

Compared with laborious pixel-wise dense labeling, it is much easier to label data by scribbles, which only costs 12 seconds to label one image. However, using scribble label…

cs.CV2021

DSC-PoseNet: Learning 6DoF Object Pose Estimation via Dual-scale Consistency

Zongxin Yang, Xin Yu, Yi Yang

Compared to 2D object bounding-box labeling, it is very difficult for humans to annotate 3D object poses, especially when depth images of scenes are unavailable. This paper investi…

cs.CV2026

ResiHMR: Residual-Limb Aware Single-Image 3D Human Mesh Recovery for Individuals with Limb Loss

Jiaying Ying, Heming Du, Kaihao Zhang +2

Single-image human mesh recovery provides a compact 3D, person-centric representation that supports analysis, animation, AR and VR, rehabilitation, and human-computer interaction.…

cs.LG2025

Train Large, Deploy Compact: Structured Compression for Compact Low-Rank Adaptation

Xin Yu, Cong Xie, Ziyu Zhao +4

Low-rank adaptation (LoRA) has become a widely used paradigm for parameter-efficient fine-tuning of large language models, yet its representational capacity often lags behind full…

cs.LG2026

Set Prediction for Next-Day Active Fire Forecasting

Yuchen Bai, Georgios Athanasiou, Xin Yu +4

Accurate next-day active fire forecasts can support early warning, disaster response, forest risk assessment, and downstream estimation of fire-related carbon emissions. Existing m…

cs.CV2026

Learning to Synergize Semantic and Geometric Priors for Limited-Data Wheat Disease Segmentation

Shijie Wang, Zijian Wang, Yadan Luo +3

Wheat disease segmentation is fundamental to precision agriculture but faces severe challenges from significant intra-class temporal variations across growth stages. Such substanti…

cs.CV2021

Write-a-speaker: Text-based Emotional and Rhythmic Talking-head Generation

Lincheng Li, Suzhen Wang, Zhimeng Zhang +4

In this paper, we propose a novel text-based talking-head video generation framework that synthesizes high-fidelity facial expressions and head motions in accordance with contextua…

eess.IV2023

Deep conditional generative models for longitudinal single-slice abdominal computed tomography harmonization

Xin Yu, Qi Yang, Yucheng Tang +8

Two-dimensional single-slice abdominal computed tomography (CT) provides a detailed tissue map with high resolution allowing quantitative characterization of relationships between…

cs.CV2025

DiPEx: Dispersing Prompt Expansion for Class-Agnostic Object Detection

Jia Syuen Lim, Zhuoxiao Chen, Mahsa Baktashmotlagh +4

Class-agnostic object detection (OD) can be a cornerstone or a bottleneck for many downstream vision tasks. Despite considerable advancements in bottom-up and multi-object discover…

cs.CV2025

Dynamic Orchestration of Multi-Agent System for Real-World Multi-Image Agricultural VQA

Yan Ke, Xin Yu, Heming Du +2

Agricultural visual question answering is essential for providing farmers and researchers with accurate and timely knowledge. However, many existing approaches are predominantly de…

cs.CV2023

The IKEA ASM Dataset: Understanding People Assembling Furniture through Actions, Objects and Pose

Yizhak Ben-Shabat, Xin Yu, Fatemeh Sadat Saleh +4

The availability of a large labeled dataset is a key requirement for applying deep learning methods to solve various computer vision tasks. In the context of understanding human ac…

cs.CV2024

TalkCLIP: Talking Head Generation with Text-Guided Expressive Speaking Styles

Yifeng Ma, Suzhen Wang, Yu Ding +6

Audio-driven talking head generation has drawn growing attention. To produce talking head videos with desired facial expressions, previous methods rely on extra reference videos to…

cs.CV2026

MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding

Fan Yang, Xingping Dong, Xin Yu +3

Understanding high-resolution (HR) images remains a critical challenge for multimodal large language models (MLLMs). Recent approaches leverage vision-based retrieval-augmented gen…

cs.CV2019

Recovering Faces from Portraits with Auxiliary Facial Attributes

Fatemeh Shiri, Xin Yu, Fatih Porikli +2

Recovering a photorealistic face from an artistic portrait is a challenging task since crucial facial details are often distorted or completely lost in artistic compositions. To ha…

physics.plasm-ph2026

Discovery of Density Limit Disruption Induced by Core-localized Alfvnic Ion Temperature Gradient Instabilities in a Tokamak Plasma

Wei Chen, Liwen Hu, Jianqiang Xu +17

To achieve a high energy gain, the fusion reactor plasma must reach a very high density. However, the tokamak plasmas ofen undergo disruption when the density exceeds the Greenwald…

cs.CV2020

High Frame Rate Video Reconstruction based on an Event Camera

Liyuan Pan, Richard Hartley, Cedric Scheerlinck +3

Event-based cameras measure intensity changes (called `events') with microsecond accuracy under high-speed motion and challenging lighting conditions. With the `active pixel sensor…

cs.CV2024

StyleTalk++: A Unified Framework for Controlling the Speaking Styles of Talking Heads

Suzhen Wang, Yifeng Ma, Yu Ding +5

Individuals have unique facial expression and head pose styles that reflect their personalized speaking styles. Existing one-shot talking head methods cannot capture such personali…

cs.CV2021

One-shot Talking Face Generation from Single-speaker Audio-Visual Correlation Learning

Suzhen Wang, Lincheng Li, Yu Ding +1

Audio-driven one-shot talking face generation methods are usually trained on video resources of various persons. However, their created videos often suffer unnatural mouth shapes a…

cs.MM2024

Diverse Sign Language Translation

Xin Shen, Lei Shen, Shaozu Yuan +3

Like spoken languages, a single sign language expression could correspond to multiple valid textual interpretations. Hence, learning a rigid one-to-one mapping for sign language tr…

nlin.PS2017

Solitons and breathers for nonisospectral mKdV equation with Darboux transformation

Ling-Jun Liu, Xin Yu

Under investigation in this paper is the nonisospectral and variable coefficients modified Kortweg-de Vries (vc-mKdV) equation, which manifests in diverse areas of physics such as…

cs.CV2020

Uncertainty-Aware Deep Calibrated Salient Object Detection

Jing Zhang, Yuchao Dai, Xin Yu +3

Existing deep neural network based salient object detection (SOD) methods mainly focus on pursuing high network accuracy. However, those methods overlook the gap between network ac…

hep-ph2014

The NLO twist-3 contributions to form factors in factorization

Shan Cheng, Ying-Ying Fan, Xin Yu +2

In this paper, we calculate the next-to-leading-order (NLO) twist-3 contribution to the form factors of transitions by employing the factorization theorem. All t…

hep-ph2013

Time-dependent CP-violations of B(Bs) decays in the perturbative QCD approach

Xin Yu, Zhi-Tian Zou, Cai-Dian Lu

We study the decay modes of B_{s}^{0}(\bar{B}_{s}^{0})-->D_{s}^{\pm} K^{\mp}, B_{s}^{0}(\bar{B}_{s}^{0})-->D^{\pm} π^{\mp} and B^{0}(\bar{B}^{0})-->D^{\pm} π^{\mp} in the perterb…

cs.CV2026

Stable Velocity: A Variance Perspective on Flow Matching

Donglin Yang, Yongxing Zhang, Xin Yu +5

While flow matching is elegant, its reliance on single-sample conditional velocities leads to high-variance training targets that destabilize optimization and slow convergence. By…

cs.CV2023

Text-Guided 3D Face Synthesis -- From Generation to Editing

Yunjie Wu, Yapeng Meng, Zhipeng Hu +5

Text-guided 3D face synthesis has achieved remarkable results by leveraging text-to-image (T2I) diffusion models. However, most existing works focus solely on the direct generation…

cs.CV2025

FingerCap: Fine-grained Finger-level Hand Motion Captioning

Xin Shen, Rui Zhu, Lei Shen +10

Understanding fine-grained human hand motion is fundamental to visual perception, embodied intelligence, and multimodal communication. In this work, we propose Fine-grained Finger-…

cs.CV2018

VLASE: Vehicle Localization by Aggregating Semantic Edges

Xin Yu, Sagar Chaturvedi, Chen Feng +4

In this paper, we propose VLASE, a framework to use semantic edge features from images to achieve on-road localization. Semantic edge features denote edge contours that separate pa…

cs.CV2023

DyGait: Exploiting Dynamic Representations for High-performance Gait Recognition

Ming Wang, Xianda Guo, Beibei Lin +5

Gait recognition is a biometric technology that recognizes the identity of humans through their walking patterns. Compared with other biometric technologies, gait recognition is mo…

cs.CV2018

Face Destylization

Fatemeh Shiri, Xin Yu, Fatih Porikli +1

Numerous style transfer methods which produce artistic styles of portraits have been proposed to date. However, the inverse problem of converting the stylized portraits back into r…

cs.CV2025

Distributed Zero-Shot Learning for Visual Recognition

Zhi Chen, Yadan Luo, Zi Huang +3

In this paper, we propose a Distributed Zero-Shot Learning (DistZSL) framework that can fully exploit decentralized data to learn an effective model for unseen classes. Considering…

cs.CV2024

Snap and Diagnose: An Advanced Multimodal Retrieval System for Identifying Plant Diseases in the Wild

Tianqi Wei, Zhi Chen, Xin Yu

Plant disease recognition is a critical task that ensures crop health and mitigates the damage caused by diseases. A handy tool that enables farmers to receive a diagnosis based on…

cs.CV2020

Mapping of Sparse 3D Data using Alternating Projection

Siddhant Ranade, Xin Yu, Shantnu Kakkar +2

We propose a novel technique to register sparse 3D scans in the absence of texture. While existing methods such as KinectFusion or Iterative Closest Points (ICP) heavily rely on de…

cs.CL2025

Decoupling Understanding from Reasoning via Problem Space Mapping for Small-Scale Model Reasoning

Li Wang, Changhao Zhang, Zengqi Xiu +4

Despite recent advances in the reasoning capabilities of Large Language Models (LLMs), improving the reasoning ability of Small Language Models (SLMs, e.g., up to 1.5B parameters)…

cs.CV2025

Cluster-Aware Prompt Ensemble Learning for Few-Shot Vision-Language Model Adaptation

Zhi Chen, Xin Yu, Xiaohui Tao +2

Vision-language models (VLMs) such as CLIP achieve zero-shot transfer across various tasks by pre-training on numerous image-text pairs. These models often benefit from using an en…