Publications (25)
Quantitative symmetry-breaking and nonlinear harmonic generation in plasmonics
Hongyu Liu, Zhi-Qiang Miao, Jingfeng Yao +2
We develop a quantitative mathematical theory that offers new perspectives on nonlinear harmonic generation in plasmonic structures arising from symmetry breaking. Focusing on seco…
From Bench to Bedside: A DeepSeek-Powered AI System for Automated Chest Radiograph Interpretation in Clinical Practice
Yaowei Bai, Ruiheng Zhang, Yu Lei +15
A global shortage of radiologists has been exacerbated by the significant volume of chest X-ray workloads, particularly in primary care. Although multimodal large language models s…
DriveLaW:Unifying Planning and Video Generation in a Latent Driving World
Tianze Xia, Yongkang Li, Lijun Zhou +9
World models have become crucial for autonomous driving, as they learn how scenarios evolve over time to address the long-tail challenges of the real world. However, current approa…
Visual Generation Tuning
Jiahao Guo, Sinan Du, Jingfeng Yao +7
Large Vision Language Models (VLMs) effectively bridge the modality gap through extensive pretraining, acquiring sophisticated visual representations aligned with language. However…
Spatiotemporal Co-reflection with Spacetime Discontinuities at Moving Interfaces
Yongge Wang, Jingfeng Yao, Chengxun Yuan +1
The control of reflection and refraction at interfaces using engineered media is central to numerous optical technologies, with negative refraction and the suppression of backscatt…
Topological Braiding and Dynamic Probing of Phase Transitions at Temporal Interfaces in Non-Hermitian Synthetic Dimensions
Yuanhang Jiang, Jianfei Li, Chengxi Yang +6
Non-Hermitian systems give rise to distinct topological phenomena, yet their manifestations at temporal interfaces characterized by abrupt changes in system parameters remain large…
MobileI2V: Fast and High-Resolution Image-to-Video on Mobile Devices
Shuai Zhang, Bao Tang, Siyuan Yu +7
Recently, video generation has witnessed rapid advancements, drawing increasing attention to image-to-video (I2V) synthesis on mobile devices. However, the substantial computationa…
Boundary conditions for drift-diffusion equations in gas-discharge plasmas
V. V. Gorin, A. A. Kudryavtsev, Jingfeng Yao +2
This paper develops a general approach to the derivation of the boundary conditions for hydrodynamic equations for charged and neutral plasma components. It includes both a well-kn…
ViTMatte: Boosting Image Matting with Pretrained Plain Vision Transformers
Jingfeng Yao, Xinggang Wang, Shusheng Yang +1
Recently, plain vision Transformers (ViTs) have shown impressive performance on various computer vision tasks, thanks to their strong modeling capacity and large-scale pretraining.…
Towards Scalable Pre-training of Visual Tokenizers for Generation
Jingfeng Yao, Yuda Song, Yucong Zhou +1
The quality of the latent space in visual tokenizers (e.g., VAEs) is crucial for modern generative models. However, the standard reconstruction-based training paradigm produces a l…
UniX: Unifying Autoregression and Diffusion for Chest X-Ray Understanding and Generation
Ruiheng Zhang, Jingfeng Yao, Huangxuan Zhao +9
Despite recent progress, medical foundation models still struggle to unify visual understanding and generation, as these tasks have inherently conflicting goals: semantic abstracti…
ReWorld: Learning Better Representations for World Action Models
Tianze Xia, Lijun Zhou, Kaixin Xiong +9
World Action Models (WAMs) model future environment evolution under action conditioning, offering a scalable paradigm for autonomous driving. However, existing approaches focus lar…
Matte Anything: Interactive Natural Image Matting with Segment Anything Models
Jingfeng Yao, Xinggang Wang, Lang Ye +1
Natural image matting algorithms aim to predict the transparency map (alpha-matte) with the trimap guidance. However, the production of trimap often requires significant labor, whi…
Breaking the Limitations of Temporal Modulation via Mixed Continuity Conditions
Yongge Wang, Jingfeng Yao, Ying Wang +2
The conventional description of time-varying media assumes that electromagnetic fields evolve according to fixed continuity conditions during parameter jumps. Here we reveal that t…
A DeepSeek-Powered AI System for Automated Chest Radiograph Interpretation in Clinical Practice
Yaowei Bai, Ruiheng Zhang, Yu Lei +20
A global shortage of radiologists has been exacerbated by the significant volume of chest X-ray workloads, particularly in primary care. Although multimodal large language models s…
Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers
Gangwei Xu, Haotong Lin, Hongcheng Luo +11
This paper presents Pixel-Perfect Depth, a monocular depth estimation model based on pixel-space diffusion generation that produces high-quality, flying-pixel-free point clouds fro…
EVA-X: A Foundation Model for General Chest X-ray Analysis with Self-supervised Learning
Jingfeng Yao, Xinggang Wang, Yuehao Song +5
The diagnosis and treatment of chest diseases play a crucial role in maintaining human health. X-ray examination has become the most common clinical examination means due to its ef…
A new approach for solving the problem of creation of inverse electron distribution function and practical recommendations for experimental searches for such media in glow discharges with hollow and flat cathodes
Chengxun Yuan, E. A. Bogdanov, A. A. Kudryavtsev +2
This paper proposes a novel approach for creating an inverse electron distribution function (EDF). Based on the obtained criteria for the formation of an inverse EDF in a non-unifo…
DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models
Lunbin Zeng, Jingfeng Yao, Bencheng Liao +3
Diffusion-based decoding has recently emerged as an appealing alternative to autoregressive (AR) generation, offering the potential to update multiple tokens in parallel and reduce…
LKCell: Efficient Cell Nuclei Instance Segmentation with Large Convolution Kernels
Ziwei Cui, Jingfeng Yao, Lunbin Zeng +3
The segmentation of cell nuclei in tissue images stained with the blood dye hematoxylin and eosin (HE) is essential for various clinical applications and analyses. Due to the c…
ViTGaze: Gaze Following with Interaction Features in Vision Transformers
Yuehao Song, Xinggang Wang, Jingfeng Yao +3
Gaze following aims to interpret human-scene interactions by predicting the person's focal point of gaze. Prevailing approaches often adopt a two-stage framework, whereby multi-mod…
FasterDiT: Towards Faster Diffusion Transformers Training without Architecture Modification
Jingfeng Yao, Wang Cheng, Wenyu Liu +1
Diffusion Transformers (DiT) have attracted significant attention in research. However, they suffer from a slow convergence rate. In this paper, we aim to accelerate DiT training w…
Turbo-VAED: Fast and Stable Transfer of Video-VAEs to Mobile Devices
Ya Zou, Jingfeng Yao, Siyuan Yu +3
There is a growing demand for deploying large generative AI models on mobile devices. For recent popular video generative models, however, the Variational AutoEncoder (VAE) represe…
Topological States Decorated by Twig Boundary in Plasma Photonic Crystals
Jianfei Li, Jingfeng Yao, Ying Wang +3
The twig edge states in graphene-like structures are viewed as the fourth states complementary to their zigzag, bearded, and armchair counterparts. In this work, we study a rod-in-…
Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models
Jingfeng Yao, Bin Yang, Xinggang Wang
Latent diffusion models with Transformer architectures excel at generating high-fidelity images. However, recent studies reveal an optimization dilemma in this two-stage design: wh…