papers

Publications (44)

cs.CV2025

Qwen2.5-VL Technical Report

Shuai Bai, Keqin Chen, Xuejing Liu +24

We introduce Qwen2.5-VL, the latest flagship model of Qwen vision-language series, which demonstrates significant advancements in both foundational capabilities and innovative func…

cs.LG2020

Deep Time-Stream Framework for Click-Through Rate Prediction by Tracking Interest Evolution

Shu-Ting Shi, Wenhao Zheng, Jun Tang +4

Click-through rate (CTR) prediction is an essential task in industrial applications such as video recommendation. Recently, deep learning models have been proposed to learn the rep…

cs.CV2025

COLA: Context-aware Language-driven Test-time Adaptation

Aiming Zhang, Tianyuan Yu, Liang Bai +5

Test-time adaptation (TTA) has gained increasing popularity due to its efficacy in addressing ``distribution shift'' issue while simultaneously protecting data privacy. However, mo…

cs.CV2018

Temporally Object-based Video Co-Segmentation

Michael Ying Yang, Matthias Reso, Jun Tang +2

In this paper, we propose an unsupervised video object co-segmentation framework based on the primary object proposals to extract the common foreground object(s) from a given video…

cond-mat.supr-con2013

A Field-directional Specific Heat Study on the Gap Structure of Overdoped Ba(FeCo)As

Gang Mu, Jun Tang, Yoichi Tanabe +7

Low-temperature specific heat is measured on the overdoped Ba(Fe_{1-x}Co_x)_2As_2 (x = 0.13) single crystal under magnetic fields along three different directions. A clear anisotro…

eess.SP2025

A Two-Stage ISAC Framework for Low-Altitude Economy Based on 5G NR Signals

Haisu Wu, Hong Ren, Cunhua Pan +5

The evolution of next-generation wireless networks has spurred the vigorous development of the low-altitude economy (LAE). To support this emerging field while remaining compatible…

quant-ph2025

Phase estimation in lossy optical interferometry without a reference beam

Jun Tang, Dong-Qing Wang, Wei Zhong +2

We investigate phase estimation in a lossy interferometer using entangled coherent states, with particular focus on a scenario where no reference beam is employed. By calculating t…

cs.IT2015

Robust Design of Transmit Waveform and Receive Filter For Colocated MIMO Radar

Wei Zhu, Jun Tang

We consider the problem of angle-robust joint transmit waveform and receive filter design for colocated Multiple-Input Multiple-Output (MIMO) radar, in the presence of signal-depen…

eess.IV2025

A dataset of primary nasopharyngeal carcinoma MRI with multi-modalities segmentation

Yin Li, Qi Chen, Kai Wang +10

Multi-modality magnetic resonance imaging(MRI) data facilitate the early diagnosis, tumor segmentation, and disease staging in the management of nasopharyngeal carcinoma (NPC). The…

cs.CV2024

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Zhibo Yang, Jun Tang, Zhaohai Li +9

Large Multimodal Models (LMMs) have demonstrated impressive performance in recognizing document images with natural language instructions. However, it remains unclear to what exten…

cs.CV2025

NTIRE 2025 XGC Quality Assessment Challenge: Methods and Results

Xiaohong Liu, Xiongkuo Min, Qiang Hu +92

This paper reports on the NTIRE 2025 XGC Quality Assessment Challenge, which will be held in conjunction with the New Trends in Image Restoration and Enhancement Workshop (NTIRE) a…

eess.SP2024

Cooperative ISAC-empowered Low-Altitude Economy

Jun Tang, Yiming Yu, Cunhua Pan +4

This paper proposes a cooperative integrated sensing and communication (ISAC) scheme for the low-altitude sensing scenario, aiming at estimating the parameters of the unmanned aeri…

cs.CL2026

Rationale-Guided Knowledge Distillation for Cross-Lingual Stance Detection

Qiuli Zhou, Jingyuan Yao, Shengeng Tang +3

Stance detection aims to identify whether a text expresses a favorable or opposing attitude toward a given target, and serves as an important task for various downstream applicatio…

cs.CV2022

Vision-Language Pre-Training for Boosting Scene Text Detectors

Sibo Song, Jianqiang Wan, Zhibo Yang +4

Recently, vision-language joint representation learning has proven to be highly effective in various scenarios. In this paper, we specifically adapt vision-language joint learning…

cs.CV2026

Contrastive On-Policy Distillation

Jiacheng Ruan, Jun Tang, Wenzhen Yuan +5

On-policy Distillation (OPD) supervises a student model on trajectories sampled from its own policy by minimizing the divergence between the output distributions of the teacher and…

cond-mat.supr-con2011

Evidence for line nodes in the energy gap of the overdoped Ba(FeCo)As from low-temperature specific heat measurements

Gang Mu, Jun Tang, Yoichi Tanabe +3

Low-temperature specific heat (SH) is measured on Ba(FeCo)As single crystals in a wide doping region under different magnetic fields. For the overdoped sample…

cs.NI2016

Topology Discovery for Linear Wireless Networks with Application to Train Backbone Inauguration

Yu Liu, Jianghua Feng, Osvaldo Simeone +4

A train backbone network consists of a sequence of nodes arranged in a linear topology. A key step that enables communication in such a network is that of topology discovery, or tr…

cs.LG2026

Multivariate Time Series Anomaly Detection via Dual-Branch Reconstruction and Autoregressive Flow-based Residual Density Estimation

Jun Liu, Ying Chen, Ziqian Lu +2

Multivariate Time Series Anomaly Detection (MTSAD) is critical for real-world monitoring scenarios such as industrial control and aerospace systems. Mainstream reconstruction-based…

physics.optics2022

High sensitivity magnetic field sensor via sandwich type PDMS resonator

Weikang Xu, Jiamin Rong, Enbo Xing +5

The sandwich structure as the core layer of PDMS resonator is proposed for single-axis magnetic sensor with high sensitivity. The small Young's modulus of flexible material corresp…

cs.CV2018

Computer-aided diagnosis of lung carcinoma using deep learning - a pilot study

Zhang Li, Zheyu Hu, Jiaolong Xu +10

Aim: Early detection and correct diagnosis of lung cancer are the most important steps in improving patient outcome. This study aims to assess which deep learning models perform be…

eess.SP2026

Subspace Fusion Sensing for Cooperative ISAC

Yining Xu, Cunhua Pan, Jun Tang +2

This paper proposes a subspace fusion sensing algorithm for cooperative integrated sensing and communication. First, we stack the received signals from access points (APs) into a t…

cs.CR2017

Privacy Loss in Apple's Implementation of Differential Privacy on MacOS 10.12

Jun Tang, Aleksandra Korolova, Xiaolong Bai +2

In June 2016, Apple announced that it will deploy differential privacy for some user data collection in order to ensure privacy of user data, even from Apple. The details of Apple'…

physics.optics2022

Compact On-Chip crystalline Resonator Integration with Etching Tapered Fiber Waveguide

Jun Yue, Jiamin Rong, Enbo Xing +5

Whispering-gallery mode crystalline resonators currently maintain the best quality factor (Q) record, however, compact on-chip packaging is still a challenge although various coupl…

eess.SP2026

Complex VAE with Heavy-Tailed Likelihood for Radar Target Detection in Sea Clutter

Ting Bai, Jun Tang, Yuxin Xu

To address the heavy-tailed, spike-prone nature of sea clutter and the scarcity of labeled target data, an unsupervised complex-valued variational autoencoder (VAE) for maritime ra…

cs.CV2024

Platypus: A Generalized Specialist Model for Reading Text in Various Forms

Peng Wang, Zhaohai Li, Jun Tang +4

Reading text from images (either natural scenes or documents) has been a long-standing research topic for decades, due to the high technical challenge and wide application range. P…

cs.LG2024

Sequential Model for Predicting Patient Adherence in Subcutaneous Immunotherapy for Allergic Rhinitis

Yin Li, Yu Xiong, Wenxin Fan +6

Objective: Subcutaneous Immunotherapy (SCIT) is the long-lasting causal treatment of allergic rhinitis (AR). How to enhance the adherence of patients to maximize the benefit of all…

cs.CV2021

ARTS: Eliminating Inconsistency between Text Detection and Recognition with Auto-Rectification Text Spotter

Humen Zhong, Jun Tang, Wenhai Wang +3

Recent approaches for end-to-end text spotting have achieved promising results. However, most of the current spotters were plagued by the inconsistency problem between text detecti…

cond-mat.supr-con2010

Superconductivity induced by doping Platinum in BaFe2As2

Xiyu Zhu, Fei Han, Gang Mu +4

By substituting Fe with the 5d-transition metal Pt in BaFe2As2, we have successfully synthesized the superconductors BaFe2-xPtxAs2. The systematic evolution of the lattice constant…

cs.AI2025

The Amazon Nova Family of Models: Technical Report and Model Card

Amazon AGI, Aaron Langford, Aayush Shah +783

We present Amazon Nova, a new generation of state-of-the-art foundation models that deliver frontier intelligence and industry-leading price performance. Amazon Nova Pro is a highl…

eess.SP2018

Fast Two-Dimensional Atomic Norm Minimization in Spectrum Estimation and Denoising

Jian Pan, Jun Tang, Yong Niu

Motivated by recent work on two dimensional (2D) harmonic component recovery via atomic norm minimization (ANM), a fast 2D direction of arrival (DOA) off-grid estimation based on A…

cs.CL2026

CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing

Zhipeng Xu, Junhao Ji, Zulong Chen +10

Large Multimodal Models (LMMs) have recently shown strong performance on Optical Character Recognition (OCR) tasks, demonstrating their promising capability in document literacy. H…

cs.SD2025

Neural personal sound zones with flexible bright zone control

Wenye Zhu, Jun Tang, Xiaofei Li

Personal sound zone (PSZ) reproduction system, which attempts to create distinct virtual acoustic scenes for different listeners at their respective positions within the same spati…

cs.IT2024

Secure MIMO Communication Relying on Movable Antennas

Jun Tang, Cunhua Pan, Yang Zhang +2

This paper considers a movable antenna (MA)-aided secure multiple-input multiple-output (MIMO) communication system consisting of a base station (BS), a legitimate information rece…

cs.CV2022

PTSEFormer: Progressive Temporal-Spatial Enhanced TransFormer Towards Video Object Detection

Han Wang, Jun Tang, Xiaodong Liu +3

Recent years have witnessed a trend of applying context frames to boost the performance of object detection as video object detection. Existing methods usually aggregate features a…

cs.CV2025

Qwen3-VL Technical Report

Shuai Bai, Yuxuan Cai, Ruizhe Chen +61

We introduce Qwen3-VL, the most capable vision-language model in the Qwen series to date, achieving superior performance across a broad range of multimodal benchmarks. It natively…

cs.LG2024

An Event-centric Framework for Predicting Crime Hotspots with Flexible Time Intervals

Jiahui Jin, Yi Hong, Guandong Xu +3

Predicting crime hotspots in a city is a complex and critical task with significant societal implications. Numerous spatiotemporal correlations and irregularities pose substantial…

cond-mat.mtrl-sci2026

Roadmap on Advancements of the FHI-aims Software Package

Joseph W. Abbott, Carlos Mera Acosta, Alaa Akkoush +203

Electronic-structure theory is the foundation of the description of materials including multiscale modeling of their properties and functions. Obviously, without sufficient accurac…

cs.LG2026

scKDGM: KAN-guided Dynamic Graph Masked Learning for Single-Cell RNA-seq Clustering

Jun Tang, Pengwei Hu, Sicong Gao +3

Single-cell RNA sequencing (scRNA-seq) clustering is essential for identifying cell types, but high dimensionality, sparsity, dropout, and technical noise hinder robust expression…

stat.ME2020

Space-Time Covariance Models on Networks with An Application on Streams

Jun Tang, Dale Zimmerman

The second-order, small-scale dependence structure of a stochastic process defined in the space-time domain is key to prediction (or kriging). While great efforts have been dedicat…

cond-mat.supr-con2013

Superconductivity induced by U-doping in the SmFeAsO system

Bo Huang, Jijun Yang, Jun Tang +7

Through partial substitution of Sm by U in SmFeAsO, a different member of the family of iron-based superconductors was successfully synthesized. X-ray diffraction measurements show…

cs.CV2026

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models

Wenwen Yu, Zhibo Yang, Jianqiang Wan +5

Visually-situated text parsing (VsTP) has recently seen notable advancements, driven by the growing demand for automated document understanding and the emergence of large language…

cs.CV2024

VL-Reader: Vision and Language Reconstructor is an Effective Scene Text Recognizer

Humen Zhong, Zhibo Yang, Zhaohai Li +4

Text recognition is an inherent integration of vision and language, encompassing the visual texture in stroke patterns and the semantic context among the character sequences. Towar…

cs.CV2026

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis

Zhipeng Xu, Zulong Chen, Qing Liu +6

Key Information Extraction (KIE) converts visually rich documents into structured data, but practical deployment remains challenging: strong performance often relies on costly on-s…

cs.CV2021

MOST: A Multi-Oriented Scene Text Detector with Localization Refinement

Minghang He, Minghui Liao, Zhibo Yang +6

Over the past few years, the field of scene text detection has progressed rapidly that modern text detectors are able to hunt text in various challenging scenarios. However, they m…