papers

Publications (25)

math.AP2026

Quantitative symmetry-breaking and nonlinear harmonic generation in plasmonics

Hongyu Liu, Zhi-Qiang Miao, Jingfeng Yao +2

We develop a quantitative mathematical theory that offers new perspectives on nonlinear harmonic generation in plasmonic structures arising from symmetry breaking. Focusing on seco…

cs.HC2025

From Bench to Bedside: A DeepSeek-Powered AI System for Automated Chest Radiograph Interpretation in Clinical Practice

Yaowei Bai, Ruiheng Zhang, Yu Lei +15

A global shortage of radiologists has been exacerbated by the significant volume of chest X-ray workloads, particularly in primary care. Although multimodal large language models s…

cs.CV2026

DriveLaW:Unifying Planning and Video Generation in a Latent Driving World

Tianze Xia, Yongkang Li, Lijun Zhou +9

World models have become crucial for autonomous driving, as they learn how scenarios evolve over time to address the long-tail challenges of the real world. However, current approa…

cs.CV2025

Visual Generation Tuning

Jiahao Guo, Sinan Du, Jingfeng Yao +7

Large Vision Language Models (VLMs) effectively bridge the modality gap through extensive pretraining, acquiring sophisticated visual representations aligned with language. However…

physics.optics2026

Spatiotemporal Co-reflection with Spacetime Discontinuities at Moving Interfaces

Yongge Wang, Jingfeng Yao, Chengxun Yuan +1

The control of reflection and refraction at interfaces using engineered media is central to numerous optical technologies, with negative refraction and the suppression of backscatt…

physics.optics2025

Topological Braiding and Dynamic Probing of Phase Transitions at Temporal Interfaces in Non-Hermitian Synthetic Dimensions

Yuanhang Jiang, Jianfei Li, Chengxi Yang +6

Non-Hermitian systems give rise to distinct topological phenomena, yet their manifestations at temporal interfaces characterized by abrupt changes in system parameters remain large…

cs.CV2025

MobileI2V: Fast and High-Resolution Image-to-Video on Mobile Devices

Shuai Zhang, Bao Tang, Siyuan Yu +7

Recently, video generation has witnessed rapid advancements, drawing increasing attention to image-to-video (I2V) synthesis on mobile devices. However, the substantial computationa…

physics.plasm-ph2020

Boundary conditions for drift-diffusion equations in gas-discharge plasmas

V. V. Gorin, A. A. Kudryavtsev, Jingfeng Yao +2

This paper develops a general approach to the derivation of the boundary conditions for hydrodynamic equations for charged and neutral plasma components. It includes both a well-kn…

cs.CV2023

ViTMatte: Boosting Image Matting with Pretrained Plain Vision Transformers

Jingfeng Yao, Xinggang Wang, Shusheng Yang +1

Recently, plain vision Transformers (ViTs) have shown impressive performance on various computer vision tasks, thanks to their strong modeling capacity and large-scale pretraining.…

cs.CV2026

Towards Scalable Pre-training of Visual Tokenizers for Generation

Jingfeng Yao, Yuda Song, Yucong Zhou +1

The quality of the latent space in visual tokenizers (e.g., VAEs) is crucial for modern generative models. However, the standard reconstruction-based training paradigm produces a l…

cs.CV2026

UniX: Unifying Autoregression and Diffusion for Chest X-Ray Understanding and Generation

Ruiheng Zhang, Jingfeng Yao, Huangxuan Zhao +9

Despite recent progress, medical foundation models still struggle to unify visual understanding and generation, as these tasks have inherently conflicting goals: semantic abstracti…

cs.CV2026

ReWorld: Learning Better Representations for World Action Models

Tianze Xia, Lijun Zhou, Kaixin Xiong +9

World Action Models (WAMs) model future environment evolution under action conditioning, offering a scalable paradigm for autonomous driving. However, existing approaches focus lar…

cs.CV2024

Matte Anything: Interactive Natural Image Matting with Segment Anything Models

Jingfeng Yao, Xinggang Wang, Lang Ye +1

Natural image matting algorithms aim to predict the transparency map (alpha-matte) with the trimap guidance. However, the production of trimap often requires significant labor, whi…

physics.optics2026

Breaking the Limitations of Temporal Modulation via Mixed Continuity Conditions

Yongge Wang, Jingfeng Yao, Ying Wang +2

The conventional description of time-varying media assumes that electromagnetic fields evolve according to fixed continuity conditions during parameter jumps. Here we reveal that t…

cs.AI2025

A DeepSeek-Powered AI System for Automated Chest Radiograph Interpretation in Clinical Practice

Yaowei Bai, Ruiheng Zhang, Yu Lei +20

A global shortage of radiologists has been exacerbated by the significant volume of chest X-ray workloads, particularly in primary care. Although multimodal large language models s…

cs.CV2025

Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers

Gangwei Xu, Haotong Lin, Hongcheng Luo +11

This paper presents Pixel-Perfect Depth, a monocular depth estimation model based on pixel-space diffusion generation that produces high-quality, flying-pixel-free point clouds fro…

cs.CV2024

EVA-X: A Foundation Model for General Chest X-ray Analysis with Self-supervised Learning

Jingfeng Yao, Xinggang Wang, Yuehao Song +5

The diagnosis and treatment of chest diseases play a crucial role in maintaining human health. X-ray examination has become the most common clinical examination means due to its ef…

physics.plasm-ph2025

A new approach for solving the problem of creation of inverse electron distribution function and practical recommendations for experimental searches for such media in glow discharges with hollow and flat cathodes

Chengxun Yuan, E. A. Bogdanov, A. A. Kudryavtsev +2

This paper proposes a novel approach for creating an inverse electron distribution function (EDF). Based on the obtained criteria for the formation of an inverse EDF in a non-unifo…

cs.CV2026

DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models

Lunbin Zeng, Jingfeng Yao, Bencheng Liao +3

Diffusion-based decoding has recently emerged as an appealing alternative to autoregressive (AR) generation, offering the potential to update multiple tokens in parallel and reduce…

eess.IV2024

LKCell: Efficient Cell Nuclei Instance Segmentation with Large Convolution Kernels

Ziwei Cui, Jingfeng Yao, Lunbin Zeng +3

The segmentation of cell nuclei in tissue images stained with the blood dye hematoxylin and eosin (HE) is essential for various clinical applications and analyses. Due to the c…

cs.CV2024

ViTGaze: Gaze Following with Interaction Features in Vision Transformers

Yuehao Song, Xinggang Wang, Jingfeng Yao +3

Gaze following aims to interpret human-scene interactions by predicting the person's focal point of gaze. Prevailing approaches often adopt a two-stage framework, whereby multi-mod…

cs.CV2024

FasterDiT: Towards Faster Diffusion Transformers Training without Architecture Modification

Jingfeng Yao, Wang Cheng, Wenyu Liu +1

Diffusion Transformers (DiT) have attracted significant attention in research. However, they suffer from a slow convergence rate. In this paper, we aim to accelerate DiT training w…

cs.CV2025

Turbo-VAED: Fast and Stable Transfer of Video-VAEs to Mobile Devices

Ya Zou, Jingfeng Yao, Siyuan Yu +3

There is a growing demand for deploying large generative AI models on mobile devices. For recent popular video generative models, however, the Variational AutoEncoder (VAE) represe…

cond-mat.mes-hall2024

Topological States Decorated by Twig Boundary in Plasma Photonic Crystals

Jianfei Li, Jingfeng Yao, Ying Wang +3

The twig edge states in graphene-like structures are viewed as the fourth states complementary to their zigzag, bearded, and armchair counterparts. In this work, we study a rod-in-…

cs.CV2025

Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models

Jingfeng Yao, Bin Yang, Xinggang Wang

Latent diffusion models with Transformer architectures excel at generating high-fidelity images. However, recent studies reveal an optimization dilemma in this two-stage design: wh…