Publications (259)
Word-level Deep Sign Language Recognition from Video: A New Large-scale Dataset and Methods Comparison
Dongxu Li, Cristian Rodriguez Opazo, Xin Yu +1
Vision-based sign language recognition aims at helping deaf people to communicate with others. However, most existing sign language datasets are limited to a small number of words.…
Virtual Width Networks
Seed, Baisheng Li, Banggu Wu +115
We introduce Virtual Width Networks (VWN), a framework that delivers the benefits of wider representations without incurring the quadratic cost of increasing the hidden size. VWN d…
TAMTRL: Teacher-Aligned Reward Reshaping for Multi-Turn Reinforcement Learning in Long-Context Compression
Li Wang, Yandong Wang, Xin Yu +3
The rapid progress of large language models (LLMs) has led to remarkable performance gains across a wide range of tasks. However, when handling long documents that exceed the model…
PlantSeg: A Large-Scale In-the-wild Dataset for Plant Disease Segmentation
Tianqi Wei, Zhi Chen, Xin Yu +3
Plant diseases pose significant threats to agriculture. It necessitates proper diagnosis and effective treatment to safeguard crop yields. To automate the diagnosis process, image…
Gait Recognition via Effective Global-Local Feature Representation and Local Temporal Aggregation
Beibei Lin, Shunli Zhang, Xin Yu
Gait recognition is one of the most important biometric technologies and has been applied in many fields. Recent gait recognition frameworks represent each gait frame by descriptor…
Semileptonic decays in the perturbative QCD approach
Wen-Fei Wang, Xin Yu, Cai-Dian Lü +1
In this paper we study the semileptonic decays of (here stands for , , or ). After evaluating the $B_c^+ \to (D_{(s…
Multi-Contrast Computed Tomography Atlas of Healthy Pancreas
Yinchi Zhou, Ho Hin Lee, Yucheng Tang +6
With the substantial diversity in population demographics, such as differences in age and body composition, the volumetric morphology of pancreas varies greatly, resulting in disti…
Joint 3D Human Shape Recovery and Pose Estimation from a Single Image with Bilayer Graph
Xin Yu, Jeroen van Baar, Siheng Chen
The ability to estimate the 3D human shape and pose from images can be useful in many contexts. Recent approaches have explored using graph convolutional networks and achieved prom…
Unsupervised Adaptation of PDE Foundation Models
Ziye Song, Zhao Wei, Xin Yu +2
Pretrained partial differential equation (PDE) foundation models can generalize across different equations, but adapting them to unseen PDE systems typically requires dense solutio…
Trust-Aware Diversion for Data-Effective Distillation
Zhuojie Wu, Yanbin Liu, Xin Shen +2
Dataset distillation compresses a large dataset into a small synthetic subset that retains essential information. Existing methods assume that all samples are perfectly labeled, li…
Copy and Paste GAN: Face Hallucination from Shaded Thumbnails
Yang Zhang, Ivor Tsang, Yawei Luo +3
Existing face hallucination methods based on convolutional neural networks (CNN) have achieved impressive performance on low-resolution (LR) faces in a normal illumination conditio…
Machine Unlearning via Null Space Calibration
Huiqiang Chen, Tianqing Zhu, Xin Yu +1
Machine unlearning aims to enable models to forget specific data instances when receiving deletion requests. Current research centres on efficient unlearning to erase the influence…
ARVo: Learning All-Range Volumetric Correspondence for Video Deblurring
Dongxu Li, Chenchen Xu, Kaihao Zhang +5
Video deblurring models exploit consecutive frames to remove blurs from camera shakes and object motions. In order to utilize neighboring sharp patches, typical methods rely mainly…
Observation estimates for a semilinear heat equation in \mathbb{R}^n
Guojie Zheng, Xin Yu
This paper studies the state observation problems for the semilinear heat equation in R^n. We derive observation estimates for the equation using the logarithmic convexity property…
Understanding the Statistical Accuracy-Communication Trade-off in Personalized Federated Learning with Minimax Guarantees
Xin Yu, Zelin He, Ying Sun +2
Personalized federated learning (PFL) offers a flexible framework for aggregating information across distributed clients with heterogeneous data. This work considers a personalized…
Deep Idempotent Network for Efficient Single Image Blind Deblurring
Yuxin Mao, Zhexiong Wan, Yuchao Dai +1
Single image blind deblurring is highly ill-posed as neither the latent sharp image nor the blur kernel is known. Even though considerable progress has been made, several major dif…
EXACT: Explicit Attribute-Guided Decoding-Time Personalization
Xin Yu, Hanwen Xing, Lingzhou Xue
Achieving personalized alignment requires adapting large language models to each user's evolving context. While decoding-time personalization offers a scalable alternative to train…
VidFace: A Full-Transformer Solver for Video FaceHallucination with Unaligned Tiny Snapshots
Yuan Gan, Yawei Luo, Xin Yu +2
In this paper, we investigate the task of hallucinating an authentic high-resolution (HR) human face from multiple low-resolution (LR) video snapshots. We propose a pure transforme…
LightMover: Generative Light Movement with Color and Intensity Controls
Gengze Zhou, Tianyu Wang, Soo Ye Kim +7
We present LightMover, a framework for controllable light manipulation in single images that leverages video diffusion priors to produce physically plausible illumination changes w…
When 3D Bounding-Box Meets SAM: Point Cloud Instance Segmentation with Weak-and-Noisy Supervision
Qingtao Yu, Heming Du, Chen Liu +1
Learning from bounding-boxes annotations has shown great potential in weakly-supervised 3D point cloud instance segmentation. However, we observed that existing methods would suffe…
Estimating the Distribution of Random Parameters in a Diffusion Equation Forward Model for a Transdermal Alcohol Biosensor
Melike Sirlanci, Susan E. Luczak, Catharine E. Fairbairn +4
We estimate the distribution of random parameters in a distributed parameter model with unbounded input and output for the transdermal transport of ethanol in humans. The model tak…
Enhancing Single-Slice Segmentation with 3D-to-2D Unpaired Scan Distillation
Xin Yu, Qi Yang, Han Liu +10
2D single-slice abdominal computed tomography (CT) enables the assessment of body habitus and organ health with low radiation exposure. However, single-slice data necessitates the…
Modeling Earth-Scale Human-Like Societies with One Billion Agents
Haoxiang Guan, Jiyan He, Liyang Fan +10
Understanding the dynamic evolution of complex social phenomena requires both high-fidelity modeling of human behavior and large-scale simulations. Traditional agent-based models (…
Benchmarking In-the-wild Multimodal Disease Recognition and A Versatile Baseline
Tianqi Wei, Zhi Chen, Zi Huang +1
Existing plant disease classification models have achieved remarkable performance in recognizing in-laboratory diseased images. However, their performance often significantly degra…
OpenSight: A Simple Open-Vocabulary Framework for LiDAR-Based Object Detection
Hu Zhang, Jianhua Xu, Tao Tang +4
Traditional LiDAR-based object detection research primarily focuses on closed-set scenarios, which falls short in complex real-world applications. Directly transferring existing 2D…
Surrogate Neural Network Model for Sensitivity Analysis and Uncertainty Quantification of the Mechanical Behavior in the Optical Lens-Barrel Assembly
Shantanu Shahane, Erman Guleryuz, Diab W Abueidda +7
Surrogate neural network-based models have been lately trained and used in a variety of science and engineering applications where the number of evaluations of a target function is…
Global existence of null-form wave equations on small asymptotically Euclidean manifolds
Chengbo Wang, Xin Yu
We prove the global existence of the small solutions to the Cauchy problem for quasilinear wave equations satisfying the null condition on , where the metric is a sma…
Observation of Kondo hybridization wave in UTe2
Xin Yu, Shuikang Yu, Zheyu Wu +9
Condensed matter systems with strong electronic correlations often manifest a variety of intertwined ordered phases of charge, spin, orbital and other degrees of freedom. As a prot…
Generalized Lyapunov criteria on finite-time stability of stochastic nonlinear systems
Xin Yu, Juliang Yin, Suiyang Khoo
This paper considers the problem of finite-time stability for stochastic nonlinear systems. A new Lyapunov theorem of stochastic finite-time stability is proposed, and an important…
Exploring Active 3D Object Detection from a Generalization Perspective
Yadan Luo, Zhuoxiao Chen, Zijian Wang +3
To alleviate the high annotation cost in LiDAR-based 3D object detection, active learning is a promising solution that learns to select only a small portion of unlabeled data to an…
Hybrid Neural Rendering for Large-Scale Scenes with Motion Blur
Peng Dai, Yinda Zhang, Xin Yu +2
Rendering novel view images is highly desirable for many applications. Despite recent progress, it remains challenging to render high-fidelity and view-consistent novel views of la…
Scaling Up 3D Kernels with Bayesian Frequency Re-parameterization for Medical Image Segmentation
Ho Hin Lee, Quan Liu, Shunxing Bao +7
With the inspiration of vision transformers, the concept of depth-wise convolution revisits to provide a large Effective Receptive Field (ERF) using Large Kernel (LK) sizes for med…
Learning Strict Identity Mappings in Deep Residual Networks
Xin Yu, Zhiding Yu, Srikumar Ramalingam
A family of super deep networks, referred to as residual networks or ResNet, achieved record-beating performance in various visual tasks such as image recognition, object detection…
Affective Behaviour Analysis via Integrating Multi-Modal Knowledge
Wei Zhang, Feng Qiu, Chen Liu +4
Affective Behavior Analysis aims to facilitate technology emotionally smart, creating a world where devices can understand and react to our emotions as humans do. To comprehensivel…
RVD: A Handheld Device-Based Fundus Video Dataset for Retinal Vessel Segmentation
MD Wahiduzzaman Khan, Hongwei Sheng, Hu Zhang +11
Retinal vessel segmentation is generally grounded in image-based datasets collected with bench-top devices. The static images naturally lose the dynamic characteristics of retina f…
CF-PRNet: Coarse-to-Fine Prototype Refining Network for Point Cloud Completion and Reconstruction
Zhi Chen, Tianqi Wei, Zecheng Zhao +6
In modern agriculture, precise monitoring of plants and fruits is crucial for tasks such as high-throughput phenotyping and automated harvesting. This paper addresses the challenge…
Autonomous Stabilization of Retinal Videos for Streamlining Assessment of Spontaneous Venous Pulsations
Hongwei Sheng, Xin Yu, Feiyu Wang +4
Spontaneous retinal Venous Pulsations (SVP) are rhythmic changes in the caliber of the central retinal vein and are observed in the optic disc region (ODR) of the retina. Its absen…
Learning Object Relation Graph and Tentative Policy for Visual Navigation
Heming Du, Xin Yu, Liang Zheng
Target-driven visual navigation aims at navigating an agent towards a given target based on the observation of the agent. In this task, it is critical to learn informative visual r…
CurveShift: Is Agent Progress Scalar? Separating Level from Shape
Hanwen Xing, Pengyun Wang, BingXu Meng +8
Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate benchmark score. These summa…
Explainable LLM Unlearning Through Reasoning
Junfeng Liao, Qizhou Wang, Shanshan Ye +3
LLM unlearning is essential for mitigating safety, copyright, and privacy concerns in pre-trained large language models (LLMs). Compared to preference alignment, it offers a more e…
EfficientDreamer: High-Fidelity and Robust 3D Creation via Orthogonal-view Diffusion Prior
Zhipeng Hu, Minda Zhao, Chaoyi Zhao +6
While image diffusion models have made significant progress in text-driven 3D content creation, they often fail to accurately capture the intended meaning of text prompts, especial…
TSPNet: Hierarchical Feature Learning via Temporal Semantic Pyramid for Sign Language Translation
Dongxu Li, Chenchen Xu, Xin Yu +4
Sign language translation (SLT) aims to interpret sign video sequences into text-based natural language sentences. Sign videos consist of continuous sequences of sign gestures with…
Texture Generation on 3D Meshes with Point-UV Diffusion
Xin Yu, Peng Dai, Wenbo Li +3
In this work, we focus on synthesizing high-quality textures on 3D meshes. We present Point-UV diffusion, a coarse-to-fine pipeline that marries the denoising diffusion model with…
Reducing Positional Variance in Cross-sectional Abdominal CT Slices with Deep Conditional Generative Models
Xin Yu, Qi Yang, Yucheng Tang +8
2D low-dose single-slice abdominal computed tomography (CT) slice enables direct measurements of body composition, which are critical to quantitatively characterizing health relati…
Optimal Feature Transport for Cross-View Image Geo-Localization
Yujiao Shi, Xin Yu, Liu Liu +2
This paper addresses the problem of cross-view image geo-localization, where the geographic location of a ground-level street-view query image is estimated by matching it against a…
Beyond Static Dialogues: Benchmarking Realistic, Heterogeneous, and Evolving Long-Term Memory
Han Zhang, Zihao Tang, Xin Yu +8
In existing memory benchmarks for Large Language Models (LLMs), the evaluated dialogue sessions often lack long-term semantic consistency, and the underlying personas tend to be fl…
CMamba: Learned Image Compression with State Space Models
Zhuojie Wu, Heming Du, Shuyun Wang +4
Learned Image Compression (LIC) has explored various architectures, such as Convolutional Neural Networks (CNNs) and transformers, in modeling image content distributions in order…
Periodic parabola solitons for the nonautonomous KP equation
Yingyou Ma, Zhiqiang Chen, Xin Yu
Kadomtsev-Petviashvili (KP) equation, who can describe different models in fluids and plasmas, has drawn investigation for its solitonic solutions with various methods. In this pap…
Generalized Strichartz Estimates on Perturbed Wave Equation and Applications on Strauss Conjecture
Xin Yu
In this paper we show a general Strichartz estimate for certain perturbed wave equation, and here we can drop the nontrapping hypothesis and handle trapping obstacles with some los…
Analytical Survey of Learning with Low-Resource Data: From Analysis to Investigation
Xiaofeng Cao, Mingwei Xu, Xin Yu +8
Learning with high-resource data has demonstrated substantial success in artificial intelligence (AI); however, the costs associated with data annotation and model training remain…
Single Slice Thigh CT Muscle Group Segmentation with Domain Adaptation and Self-Training
Qi Yang, Xin Yu, Ho Hin Lee +8
Objective: Thigh muscle group segmentation is important for assessment of muscle anatomy, metabolic disease and aging. Many efforts have been put into quantifying muscle tissues wi…
Learning Implicit Body Representations from Double Diffusion Based Neural Radiance Fields
Guangming Yao, Hongzhi Wu, Yi Yuan +3
In this paper, we present a novel double diffusion based neural radiance field, dubbed DD-NeRF, to reconstruct human body geometry and render the human body appearance in novel vie…
RichRAG: Crafting Rich Responses for Multi-faceted Queries in Retrieval-Augmented Generation
Shuting Wang, Xin Yu, Mang Wang +3
Retrieval-augmented generation (RAG) effectively addresses issues of static knowledge and hallucination in large language models. Existing studies mostly focus on question scenario…
Scaling Up Exact Neural Network Compression by ReLU Stability
Thiago Serra, Xin Yu, Abhinav Kumar +1
We can compress a rectifier network while exactly preserving its underlying functionality with respect to a given input domain if some of its neurons are stable. However, current a…
UniTok: A Unified Tokenizer for Visual Generation and Understanding
Chuofan Ma, Yi Jiang, Junfeng Wu +5
Visual generative and understanding models typically rely on distinct tokenizers to process images, presenting a key challenge for unifying them within a single framework. Recent s…
Enhancing Hierarchical Transformers for Whole Brain Segmentation with Intracranial Measurements Integration
Xin Yu, Yucheng Tang, Qi Yang +4
Whole brain segmentation with magnetic resonance imaging (MRI) enables the non-invasive measurement of brain regions, including total intracranial volume (TICV) and posterior fossa…
Effective Graph Resistance as Cumulative Heat Dissipation
Xiangrong Wang, Xin Yu, Zongze Wu +1
Effective graph resistance is a fundamental structural metric in network science, widely used to quantify global connectivity, compare network architectures, and assess robustness…
NVIDIA Nemotron Parse 1.1
Kateryna Chumachenko, Amala Sanjay Deshmukh, Jarno Seppanen +30
We introduce Nemotron-Parse-1.1, a lightweight document parsing and OCR model that advances the capabilities of its predecessor, Nemoretriever-Parse-1.0. Nemotron-Parse-1.1 deliver…
Instance As Identity: A Generic Online Paradigm for Video Instance Segmentation
Feng Zhu, Zongxin Yang, Xin Yu +2
Modeling temporal information for both detection and tracking in a unified framework has been proved a promising solution to video instance segmentation (VIS). However, how to effe…
The Combinatorial Brain Surgeon: Pruning Weights That Cancel One Another in Neural Networks
Xin Yu, Thiago Serra, Srikumar Ramalingam +1
Neural networks tend to achieve better accuracy with training if they are larger -- even if the resulting models are overparameterized. Nevertheless, carefully removing such excess…
Concerning the Strauss conjecture on asymptotically Euclidean manifolds
Chengbo Wang, Xin Yu
In this paper we verify the Strauss conjecture for semilinear wave equations on asymptotically Euclidean manifolds when n=3,4, we also give an almost sharp life span for the subcri…
Where am I looking at? Joint Location and Orientation Estimation by Cross-View Matching
Yujiao Shi, Xin Yu, Dylan Campbell +1
Cross-view geo-localization is the problem of estimating the position and orientation (latitude, longitude and azimuth angle) of a camera at ground level given a large-scale databa…
Charmed Scalar Meson Production in Decays
Yue-Long Shen, Xin Yu
The study on the charmed scalar meson spectroscopy has become a hot topic both experimentally and theoretically. The decays provide an ideal place to study their property…
Generalized and weighted Strichartz estimates
Jin-Cheng Jiang, Chengbo Wang, Xin Yu
In this paper, we explore the relations between different kinds of Strichartz estimates and give new estimates in Euclidean space . In particular, we prove the genera…
The Sound of Risk: A Multimodal Physics-Informed Acoustic Model for Forecasting Market Volatility and Enhancing Market Interpretability
Xiaoliang Chen, Xin Yu, Le Chang +8
Information asymmetry in financial markets, often amplified by strategically crafted corporate narratives, undermines the effectiveness of conventional textual analysis. We propose…
Heterogeneity in Multi-Agent Reinforcement Learning
Tianyi Hu, Zhiqiang Pu, Yuan Wang +3
Heterogeneity is a fundamental property in multi-agent reinforcement learning (MARL), which is closely related not only to the functional differences of agents, but also to policy…
Weakly-Supervised Salient Object Detection via Scribble Annotations
Jing Zhang, Xin Yu, Aixuan Li +3
Compared with laborious pixel-wise dense labeling, it is much easier to label data by scribbles, which only costs 12 seconds to label one image. However, using scribble label…
DSC-PoseNet: Learning 6DoF Object Pose Estimation via Dual-scale Consistency
Zongxin Yang, Xin Yu, Yi Yang
Compared to 2D object bounding-box labeling, it is very difficult for humans to annotate 3D object poses, especially when depth images of scenes are unavailable. This paper investi…
ResiHMR: Residual-Limb Aware Single-Image 3D Human Mesh Recovery for Individuals with Limb Loss
Jiaying Ying, Heming Du, Kaihao Zhang +2
Single-image human mesh recovery provides a compact 3D, person-centric representation that supports analysis, animation, AR and VR, rehabilitation, and human-computer interaction.…
Train Large, Deploy Compact: Structured Compression for Compact Low-Rank Adaptation
Xin Yu, Cong Xie, Ziyu Zhao +4
Low-rank adaptation (LoRA) has become a widely used paradigm for parameter-efficient fine-tuning of large language models, yet its representational capacity often lags behind full…
Set Prediction for Next-Day Active Fire Forecasting
Yuchen Bai, Georgios Athanasiou, Xin Yu +4
Accurate next-day active fire forecasts can support early warning, disaster response, forest risk assessment, and downstream estimation of fire-related carbon emissions. Existing m…
Learning to Synergize Semantic and Geometric Priors for Limited-Data Wheat Disease Segmentation
Shijie Wang, Zijian Wang, Yadan Luo +3
Wheat disease segmentation is fundamental to precision agriculture but faces severe challenges from significant intra-class temporal variations across growth stages. Such substanti…
Write-a-speaker: Text-based Emotional and Rhythmic Talking-head Generation
Lincheng Li, Suzhen Wang, Zhimeng Zhang +4
In this paper, we propose a novel text-based talking-head video generation framework that synthesizes high-fidelity facial expressions and head motions in accordance with contextua…
Deep conditional generative models for longitudinal single-slice abdominal computed tomography harmonization
Xin Yu, Qi Yang, Yucheng Tang +8
Two-dimensional single-slice abdominal computed tomography (CT) provides a detailed tissue map with high resolution allowing quantitative characterization of relationships between…
DiPEx: Dispersing Prompt Expansion for Class-Agnostic Object Detection
Jia Syuen Lim, Zhuoxiao Chen, Mahsa Baktashmotlagh +4
Class-agnostic object detection (OD) can be a cornerstone or a bottleneck for many downstream vision tasks. Despite considerable advancements in bottom-up and multi-object discover…
Dynamic Orchestration of Multi-Agent System for Real-World Multi-Image Agricultural VQA
Yan Ke, Xin Yu, Heming Du +2
Agricultural visual question answering is essential for providing farmers and researchers with accurate and timely knowledge. However, many existing approaches are predominantly de…
The IKEA ASM Dataset: Understanding People Assembling Furniture through Actions, Objects and Pose
Yizhak Ben-Shabat, Xin Yu, Fatemeh Sadat Saleh +4
The availability of a large labeled dataset is a key requirement for applying deep learning methods to solve various computer vision tasks. In the context of understanding human ac…
TalkCLIP: Talking Head Generation with Text-Guided Expressive Speaking Styles
Yifeng Ma, Suzhen Wang, Yu Ding +6
Audio-driven talking head generation has drawn growing attention. To produce talking head videos with desired facial expressions, previous methods rely on extra reference videos to…
MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding
Fan Yang, Xingping Dong, Xin Yu +3
Understanding high-resolution (HR) images remains a critical challenge for multimodal large language models (MLLMs). Recent approaches leverage vision-based retrieval-augmented gen…
Recovering Faces from Portraits with Auxiliary Facial Attributes
Fatemeh Shiri, Xin Yu, Fatih Porikli +2
Recovering a photorealistic face from an artistic portrait is a challenging task since crucial facial details are often distorted or completely lost in artistic compositions. To ha…
Discovery of Density Limit Disruption Induced by Core-localized Alfvnic Ion Temperature Gradient Instabilities in a Tokamak Plasma
Wei Chen, Liwen Hu, Jianqiang Xu +17
To achieve a high energy gain, the fusion reactor plasma must reach a very high density. However, the tokamak plasmas ofen undergo disruption when the density exceeds the Greenwald…
High Frame Rate Video Reconstruction based on an Event Camera
Liyuan Pan, Richard Hartley, Cedric Scheerlinck +3
Event-based cameras measure intensity changes (called `events') with microsecond accuracy under high-speed motion and challenging lighting conditions. With the `active pixel sensor…
StyleTalk++: A Unified Framework for Controlling the Speaking Styles of Talking Heads
Suzhen Wang, Yifeng Ma, Yu Ding +5
Individuals have unique facial expression and head pose styles that reflect their personalized speaking styles. Existing one-shot talking head methods cannot capture such personali…
One-shot Talking Face Generation from Single-speaker Audio-Visual Correlation Learning
Suzhen Wang, Lincheng Li, Yu Ding +1
Audio-driven one-shot talking face generation methods are usually trained on video resources of various persons. However, their created videos often suffer unnatural mouth shapes a…
Diverse Sign Language Translation
Xin Shen, Lei Shen, Shaozu Yuan +3
Like spoken languages, a single sign language expression could correspond to multiple valid textual interpretations. Hence, learning a rigid one-to-one mapping for sign language tr…
Solitons and breathers for nonisospectral mKdV equation with Darboux transformation
Ling-Jun Liu, Xin Yu
Under investigation in this paper is the nonisospectral and variable coefficients modified Kortweg-de Vries (vc-mKdV) equation, which manifests in diverse areas of physics such as…
Uncertainty-Aware Deep Calibrated Salient Object Detection
Jing Zhang, Yuchao Dai, Xin Yu +3
Existing deep neural network based salient object detection (SOD) methods mainly focus on pursuing high network accuracy. However, those methods overlook the gap between network ac…
The NLO twist-3 contributions to form factors in factorization
Shan Cheng, Ying-Ying Fan, Xin Yu +2
In this paper, we calculate the next-to-leading-order (NLO) twist-3 contribution to the form factors of transitions by employing the factorization theorem. All t…
Time-dependent CP-violations of B(Bs) decays in the perturbative QCD approach
Xin Yu, Zhi-Tian Zou, Cai-Dian Lu
We study the decay modes of B_{s}^{0}(\bar{B}_{s}^{0})-->D_{s}^{\pm} K^{\mp}, B_{s}^{0}(\bar{B}_{s}^{0})-->D^{\pm} Ï^{\mp} and B^{0}(\bar{B}^{0})-->D^{\pm} Ï^{\mp} in the perterb…
Stable Velocity: A Variance Perspective on Flow Matching
Donglin Yang, Yongxing Zhang, Xin Yu +5
While flow matching is elegant, its reliance on single-sample conditional velocities leads to high-variance training targets that destabilize optimization and slow convergence. By…
Text-Guided 3D Face Synthesis -- From Generation to Editing
Yunjie Wu, Yapeng Meng, Zhipeng Hu +5
Text-guided 3D face synthesis has achieved remarkable results by leveraging text-to-image (T2I) diffusion models. However, most existing works focus solely on the direct generation…
FingerCap: Fine-grained Finger-level Hand Motion Captioning
Xin Shen, Rui Zhu, Lei Shen +10
Understanding fine-grained human hand motion is fundamental to visual perception, embodied intelligence, and multimodal communication. In this work, we propose Fine-grained Finger-…
VLASE: Vehicle Localization by Aggregating Semantic Edges
Xin Yu, Sagar Chaturvedi, Chen Feng +4
In this paper, we propose VLASE, a framework to use semantic edge features from images to achieve on-road localization. Semantic edge features denote edge contours that separate pa…
DyGait: Exploiting Dynamic Representations for High-performance Gait Recognition
Ming Wang, Xianda Guo, Beibei Lin +5
Gait recognition is a biometric technology that recognizes the identity of humans through their walking patterns. Compared with other biometric technologies, gait recognition is mo…
Face Destylization
Fatemeh Shiri, Xin Yu, Fatih Porikli +1
Numerous style transfer methods which produce artistic styles of portraits have been proposed to date. However, the inverse problem of converting the stylized portraits back into r…
Distributed Zero-Shot Learning for Visual Recognition
Zhi Chen, Yadan Luo, Zi Huang +3
In this paper, we propose a Distributed Zero-Shot Learning (DistZSL) framework that can fully exploit decentralized data to learn an effective model for unseen classes. Considering…
Snap and Diagnose: An Advanced Multimodal Retrieval System for Identifying Plant Diseases in the Wild
Tianqi Wei, Zhi Chen, Xin Yu
Plant disease recognition is a critical task that ensures crop health and mitigates the damage caused by diseases. A handy tool that enables farmers to receive a diagnosis based on…
Mapping of Sparse 3D Data using Alternating Projection
Siddhant Ranade, Xin Yu, Shantnu Kakkar +2
We propose a novel technique to register sparse 3D scans in the absence of texture. While existing methods such as KinectFusion or Iterative Closest Points (ICP) heavily rely on de…
Decoupling Understanding from Reasoning via Problem Space Mapping for Small-Scale Model Reasoning
Li Wang, Changhao Zhang, Zengqi Xiu +4
Despite recent advances in the reasoning capabilities of Large Language Models (LLMs), improving the reasoning ability of Small Language Models (SLMs, e.g., up to 1.5B parameters)…
Cluster-Aware Prompt Ensemble Learning for Few-Shot Vision-Language Model Adaptation
Zhi Chen, Xin Yu, Xiaohui Tao +2
Vision-language models (VLMs) such as CLIP achieve zero-shot transfer across various tasks by pre-training on numerous image-text pairs. These models often benefit from using an en…