Publications (44)
Qwen2.5-VL Technical Report
Shuai Bai, Keqin Chen, Xuejing Liu +24
We introduce Qwen2.5-VL, the latest flagship model of Qwen vision-language series, which demonstrates significant advancements in both foundational capabilities and innovative func…
Deep Time-Stream Framework for Click-Through Rate Prediction by Tracking Interest Evolution
Shu-Ting Shi, Wenhao Zheng, Jun Tang +4
Click-through rate (CTR) prediction is an essential task in industrial applications such as video recommendation. Recently, deep learning models have been proposed to learn the rep…
COLA: Context-aware Language-driven Test-time Adaptation
Aiming Zhang, Tianyuan Yu, Liang Bai +5
Test-time adaptation (TTA) has gained increasing popularity due to its efficacy in addressing ``distribution shift'' issue while simultaneously protecting data privacy. However, mo…
Temporally Object-based Video Co-Segmentation
Michael Ying Yang, Matthias Reso, Jun Tang +2
In this paper, we propose an unsupervised video object co-segmentation framework based on the primary object proposals to extract the common foreground object(s) from a given video…
A Field-directional Specific Heat Study on the Gap Structure of Overdoped Ba(FeCo)As
Gang Mu, Jun Tang, Yoichi Tanabe +7
Low-temperature specific heat is measured on the overdoped Ba(Fe_{1-x}Co_x)_2As_2 (x = 0.13) single crystal under magnetic fields along three different directions. A clear anisotro…
A Two-Stage ISAC Framework for Low-Altitude Economy Based on 5G NR Signals
Haisu Wu, Hong Ren, Cunhua Pan +5
The evolution of next-generation wireless networks has spurred the vigorous development of the low-altitude economy (LAE). To support this emerging field while remaining compatible…
Phase estimation in lossy optical interferometry without a reference beam
Jun Tang, Dong-Qing Wang, Wei Zhong +2
We investigate phase estimation in a lossy interferometer using entangled coherent states, with particular focus on a scenario where no reference beam is employed. By calculating t…
Robust Design of Transmit Waveform and Receive Filter For Colocated MIMO Radar
Wei Zhu, Jun Tang
We consider the problem of angle-robust joint transmit waveform and receive filter design for colocated Multiple-Input Multiple-Output (MIMO) radar, in the presence of signal-depen…
A dataset of primary nasopharyngeal carcinoma MRI with multi-modalities segmentation
Yin Li, Qi Chen, Kai Wang +10
Multi-modality magnetic resonance imaging(MRI) data facilitate the early diagnosis, tumor segmentation, and disease staging in the management of nasopharyngeal carcinoma (NPC). The…
CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy
Zhibo Yang, Jun Tang, Zhaohai Li +9
Large Multimodal Models (LMMs) have demonstrated impressive performance in recognizing document images with natural language instructions. However, it remains unclear to what exten…
NTIRE 2025 XGC Quality Assessment Challenge: Methods and Results
Xiaohong Liu, Xiongkuo Min, Qiang Hu +92
This paper reports on the NTIRE 2025 XGC Quality Assessment Challenge, which will be held in conjunction with the New Trends in Image Restoration and Enhancement Workshop (NTIRE) a…
Cooperative ISAC-empowered Low-Altitude Economy
Jun Tang, Yiming Yu, Cunhua Pan +4
This paper proposes a cooperative integrated sensing and communication (ISAC) scheme for the low-altitude sensing scenario, aiming at estimating the parameters of the unmanned aeri…
Rationale-Guided Knowledge Distillation for Cross-Lingual Stance Detection
Qiuli Zhou, Jingyuan Yao, Shengeng Tang +3
Stance detection aims to identify whether a text expresses a favorable or opposing attitude toward a given target, and serves as an important task for various downstream applicatio…
Vision-Language Pre-Training for Boosting Scene Text Detectors
Sibo Song, Jianqiang Wan, Zhibo Yang +4
Recently, vision-language joint representation learning has proven to be highly effective in various scenarios. In this paper, we specifically adapt vision-language joint learning…
Contrastive On-Policy Distillation
Jiacheng Ruan, Jun Tang, Wenzhen Yuan +5
On-policy Distillation (OPD) supervises a student model on trajectories sampled from its own policy by minimizing the divergence between the output distributions of the teacher and…
Evidence for line nodes in the energy gap of the overdoped Ba(FeCo)As from low-temperature specific heat measurements
Gang Mu, Jun Tang, Yoichi Tanabe +3
Low-temperature specific heat (SH) is measured on Ba(FeCo)As single crystals in a wide doping region under different magnetic fields. For the overdoped sample…
Topology Discovery for Linear Wireless Networks with Application to Train Backbone Inauguration
Yu Liu, Jianghua Feng, Osvaldo Simeone +4
A train backbone network consists of a sequence of nodes arranged in a linear topology. A key step that enables communication in such a network is that of topology discovery, or tr…
Multivariate Time Series Anomaly Detection via Dual-Branch Reconstruction and Autoregressive Flow-based Residual Density Estimation
Jun Liu, Ying Chen, Ziqian Lu +2
Multivariate Time Series Anomaly Detection (MTSAD) is critical for real-world monitoring scenarios such as industrial control and aerospace systems. Mainstream reconstruction-based…
High sensitivity magnetic field sensor via sandwich type PDMS resonator
Weikang Xu, Jiamin Rong, Enbo Xing +5
The sandwich structure as the core layer of PDMS resonator is proposed for single-axis magnetic sensor with high sensitivity. The small Young's modulus of flexible material corresp…
Computer-aided diagnosis of lung carcinoma using deep learning - a pilot study
Zhang Li, Zheyu Hu, Jiaolong Xu +10
Aim: Early detection and correct diagnosis of lung cancer are the most important steps in improving patient outcome. This study aims to assess which deep learning models perform be…
Subspace Fusion Sensing for Cooperative ISAC
Yining Xu, Cunhua Pan, Jun Tang +2
This paper proposes a subspace fusion sensing algorithm for cooperative integrated sensing and communication. First, we stack the received signals from access points (APs) into a t…
Privacy Loss in Apple's Implementation of Differential Privacy on MacOS 10.12
Jun Tang, Aleksandra Korolova, Xiaolong Bai +2
In June 2016, Apple announced that it will deploy differential privacy for some user data collection in order to ensure privacy of user data, even from Apple. The details of Apple'…
Compact On-Chip crystalline Resonator Integration with Etching Tapered Fiber Waveguide
Jun Yue, Jiamin Rong, Enbo Xing +5
Whispering-gallery mode crystalline resonators currently maintain the best quality factor (Q) record, however, compact on-chip packaging is still a challenge although various coupl…
Complex VAE with Heavy-Tailed Likelihood for Radar Target Detection in Sea Clutter
Ting Bai, Jun Tang, Yuxin Xu
To address the heavy-tailed, spike-prone nature of sea clutter and the scarcity of labeled target data, an unsupervised complex-valued variational autoencoder (VAE) for maritime ra…
Platypus: A Generalized Specialist Model for Reading Text in Various Forms
Peng Wang, Zhaohai Li, Jun Tang +4
Reading text from images (either natural scenes or documents) has been a long-standing research topic for decades, due to the high technical challenge and wide application range. P…
Sequential Model for Predicting Patient Adherence in Subcutaneous Immunotherapy for Allergic Rhinitis
Yin Li, Yu Xiong, Wenxin Fan +6
Objective: Subcutaneous Immunotherapy (SCIT) is the long-lasting causal treatment of allergic rhinitis (AR). How to enhance the adherence of patients to maximize the benefit of all…
ARTS: Eliminating Inconsistency between Text Detection and Recognition with Auto-Rectification Text Spotter
Humen Zhong, Jun Tang, Wenhai Wang +3
Recent approaches for end-to-end text spotting have achieved promising results. However, most of the current spotters were plagued by the inconsistency problem between text detecti…
Superconductivity induced by doping Platinum in BaFe2As2
Xiyu Zhu, Fei Han, Gang Mu +4
By substituting Fe with the 5d-transition metal Pt in BaFe2As2, we have successfully synthesized the superconductors BaFe2-xPtxAs2. The systematic evolution of the lattice constant…
The Amazon Nova Family of Models: Technical Report and Model Card
Amazon AGI, Aaron Langford, Aayush Shah +783
We present Amazon Nova, a new generation of state-of-the-art foundation models that deliver frontier intelligence and industry-leading price performance. Amazon Nova Pro is a highl…
Fast Two-Dimensional Atomic Norm Minimization in Spectrum Estimation and Denoising
Jian Pan, Jun Tang, Yong Niu
Motivated by recent work on two dimensional (2D) harmonic component recovery via atomic norm minimization (ANM), a fast 2D direction of arrival (DOA) off-grid estimation based on A…
CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing
Zhipeng Xu, Junhao Ji, Zulong Chen +10
Large Multimodal Models (LMMs) have recently shown strong performance on Optical Character Recognition (OCR) tasks, demonstrating their promising capability in document literacy. H…
Neural personal sound zones with flexible bright zone control
Wenye Zhu, Jun Tang, Xiaofei Li
Personal sound zone (PSZ) reproduction system, which attempts to create distinct virtual acoustic scenes for different listeners at their respective positions within the same spati…
Secure MIMO Communication Relying on Movable Antennas
Jun Tang, Cunhua Pan, Yang Zhang +2
This paper considers a movable antenna (MA)-aided secure multiple-input multiple-output (MIMO) communication system consisting of a base station (BS), a legitimate information rece…
PTSEFormer: Progressive Temporal-Spatial Enhanced TransFormer Towards Video Object Detection
Han Wang, Jun Tang, Xiaodong Liu +3
Recent years have witnessed a trend of applying context frames to boost the performance of object detection as video object detection. Existing methods usually aggregate features a…
Qwen3-VL Technical Report
Shuai Bai, Yuxuan Cai, Ruizhe Chen +61
We introduce Qwen3-VL, the most capable vision-language model in the Qwen series to date, achieving superior performance across a broad range of multimodal benchmarks. It natively…
An Event-centric Framework for Predicting Crime Hotspots with Flexible Time Intervals
Jiahui Jin, Yi Hong, Guandong Xu +3
Predicting crime hotspots in a city is a complex and critical task with significant societal implications. Numerous spatiotemporal correlations and irregularities pose substantial…
Roadmap on Advancements of the FHI-aims Software Package
Joseph W. Abbott, Carlos Mera Acosta, Alaa Akkoush +203
Electronic-structure theory is the foundation of the description of materials including multiscale modeling of their properties and functions. Obviously, without sufficient accurac…
scKDGM: KAN-guided Dynamic Graph Masked Learning for Single-Cell RNA-seq Clustering
Jun Tang, Pengwei Hu, Sicong Gao +3
Single-cell RNA sequencing (scRNA-seq) clustering is essential for identifying cell types, but high dimensionality, sparsity, dropout, and technical noise hinder robust expression…
Space-Time Covariance Models on Networks with An Application on Streams
Jun Tang, Dale Zimmerman
The second-order, small-scale dependence structure of a stochastic process defined in the space-time domain is key to prediction (or kriging). While great efforts have been dedicat…
Superconductivity induced by U-doping in the SmFeAsO system
Bo Huang, Jijun Yang, Jun Tang +7
Through partial substitution of Sm by U in SmFeAsO, a different member of the family of iron-based superconductors was successfully synthesized. X-ray diffraction measurements show…
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models
Wenwen Yu, Zhibo Yang, Jianqiang Wan +5
Visually-situated text parsing (VsTP) has recently seen notable advancements, driven by the growing demand for automated document understanding and the emergence of large language…
VL-Reader: Vision and Language Reconstructor is an Effective Scene Text Recognizer
Humen Zhong, Zhibo Yang, Zhaohai Li +4
Text recognition is an inherent integration of vision and language, encompassing the visual texture in stroke patterns and the semantic context among the character sequences. Towar…
Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis
Zhipeng Xu, Zulong Chen, Qing Liu +6
Key Information Extraction (KIE) converts visually rich documents into structured data, but practical deployment remains challenging: strong performance often relies on costly on-s…
MOST: A Multi-Oriented Scene Text Detector with Localization Refinement
Minghang He, Minghui Liao, Zhibo Yang +6
Over the past few years, the field of scene text detection has progressed rapidly that modern text detectors are able to hunt text in various challenging scenarios. However, they m…