papers

Publications (153)

cs.CL2026

Controllable Spoken Dialogue Generation: An LLM-Driven Grading System for K-12 Non-Native English Learners

Haidong Yuan, Haokun Zhao, Wanshi Xu +4

Large language models (LLMs) often fail to meet the pedagogical needs of K-12 English learners in non-native contexts due to a proficiency mismatch. To address this widespread chal…

cond-mat.str-el2024

Low-energy spin dynamics in a Kitaev material Na3Ni2BiO6 investigated by NMR

Xinyu Shi, Yi Cui, Yanyan Shangguan +11

We performed 23Na NMR and magnetization measurements on an S = 1, quasi-2D honeycomb lattice antiferromagnet Na3Ni2BiO6. A large positive Curie-Weiss constant of 22.9 K is observed…

cs.CV2025

Infrared and Visible Image Fusion: From Data Compatibility to Task Adaption

Jinyuan Liu, Guanyao Wu, Zhu Liu +6

Infrared-visible image fusion (IVIF) is a critical task in computer vision, aimed at integrating the unique features of both infrared and visible spectra into a unified representat…

cs.AI2024

Fast Peer Adaptation with Context-aware Exploration

Long Ma, Yuanfei Wang, Fangwei Zhong +2

Fast adapting to unknown peers (partners or opponents) with different strategies is a key challenge in multi-agent games. To do so, it is crucial for the agent to probe and identif…

eess.AS2021

Improving Hybrid CTC/Attention End-to-end Speech Recognition with Pretrained Acoustic and Language Model

Keqi Deng, Songjun Cao, Yike Zhang +1

Recently, self-supervised pretraining has achieved impressive results in end-to-end (E2E) automatic speech recognition (ASR). However, the dominant sequence-to-sequence (S2S) E2E m…

eess.AS2021

Improving Accent Identification and Accented Speech Recognition Under a Framework of Self-supervised Learning

Keqi Deng, Songjun Cao, Long Ma

Recently, self-supervised pre-training has gained success in automatic speech recognition (ASR). However, considering the difference between speech accents in real scenarios, how t…

eess.AS2023

DistillW2V2: A Small and Streaming Wav2vec 2.0 Based ASR Model

Yanzhe Fu, Yueteng Kang, Songjun Cao +1

Wav2vec 2.0 (W2V2) has shown impressive performance in automatic speech recognition (ASR). However, the large model size and the non-streaming architecture make it hard to be used…

cs.CL2026

Can Deep Research Agents Retrieve and Organize? Evaluating the Synthesis Gap with Expert Taxonomies

Ming Zhang, Jiabao Zhuang, Wenqing Jing +18

Deep Research Agents increasingly automate survey generation, yet whether they match human experts at retrieving essential papers and organizing them into expert-like taxonomies re…

physics.chem-ph2021

Significantly Enhanced Performance of Nanofluidic Osmotic Power Generation by Slipping Surfaces of Nanopores

Long Ma, Kabin Lin, Yinghua Qiu +4

High-performance osmotic energy conversion (OEC) with perm-selective porous membrane requires both high ionic selectivity and permeability simultaneously. Here, hydrodynamic slip i…

cs.AI2026

Robust and Generalizable Safety Steering for Text-to-Image Diffusion Transformers

Zihao Xue, Yan Wang, Zhen Bi +7

Diffusion Transformers have become a powerful backbone for text-to-image generation, but their layered and cross-modal generation process makes safety control fundamentally differe…

cs.LG2026

Route Experts by Sequence, not by Token

Tiansheng Wen, Yifei Wang, Aosong Feng +7

Mixture-of-Experts (MoE) architectures scale large language models (LLMs) by activating only a subset of experts per token, but the standard TopK routing assigns the same fixed num…

cs.CL2022

Improving CTC-based speech recognition via knowledge transferring from pre-trained language models

Keqi Deng, Songjun Cao, Yike Zhang +4

Recently, end-to-end automatic speech recognition models based on connectionist temporal classification (CTC) have achieved impressive results, especially when fine-tuned from wav2…

cs.CV2026

Learning with Semantic Priors: Stabilizing Point-Supervised Infrared Small Target Detection via Hierarchical Knowledge Distillation

Yuanhang Yao, Ping Qian, Zhu Liu +2

Single-frame Infrared Small Target Detection (ISTD) aims to localize weak targets under heavy background clutter, yet dense pixel-wise annotations are expensive. Point supervision…

cond-mat.supr-con2025

Preformed Cooper Pairs in a Triclinic Iron Pnictide Superconductor

Zezhong Li, Wenshan Hong, Honglin Zhou +10

Electron pairing along with phase coherence generates superconductivity below the critical temperature (). In underdoped high- cuprates, these two quantum phenomena may o…

cs.CV2023

Hybrid-Supervised Dual-Search: Leveraging Automatic Learning for Loss-free Multi-Exposure Image Fusion

Guanyao Wu, Hongming Fu, Jinyuan Liu +3

Multi-exposure image fusion (MEF) has emerged as a prominent solution to address the limitations of digital imaging in representing varied exposure levels. Despite its advancements…

cs.CV2023

Bilevel Fast Scene Adaptation for Low-Light Image Enhancement

Long Ma, Dian Jin, Nan An +3

Enhancing images in low-light scenes is a challenging but widely concerned task in the computer vision. The mainstream learning-based methods mainly acquire the enhanced model by l…

cs.CV2023

Enhancing Infrared Small Target Detection Robustness with Bi-Level Adversarial Framework

Zhu Liu, Zihang Chen, Jinyuan Liu +3

The detection of small infrared targets against blurred and cluttered backgrounds has remained an enduring challenge. In recent years, learning-based schemes have become the mainst…

physics.chem-ph2022

Effective Charged Exterior Surfaces for Enhanced Ionic Diffusion through Nanopores under Salt Gradients

Long Ma, Xuan An, Fenhong Song +1

High-performance osmotic energy conversion requires both large ionic throughput and high ionic selectivity, which can be significantly promoted by exterior surface charges simultan…

cond-mat.str-el2019

Novel magnetic field tuning of quantum spin excitations in a weakly coupled 1/2 Heisenberg spin chain as seen from NMR

Long Ma, Z. Wang, L. Hu +3

We report our NMR study of the spin excitations in the quasi-one dimensional (1D) quantum magnet CHNHCu(HCOO) lying in the 1D-3D dimensional crossover regime wi…

cs.SD2025

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis

Rui Niu, Weihao Wu, Jie Chen +2

Controllable speech synthesis aims to control the style of generated speech using reference input, which can be of various modalities. Existing face-based methods struggle with rob…

cs.SD2024

Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Xiong Wang, Yangze Li, Chaoyou Fu +5

Rapidly developing large language models (LLMs) have brought tremendous intelligent applications. Especially, the GPT-4o's excellent duplex speech interaction ability has brought i…

q-bio.GN2013

Methods for scoring the collective effect of SNPs: Minor alleles of common SNPs quantitatively affect traits/diseases and are under both positive and negative selection

Dejian Yuan, Zuobin Zhu, Xiaohua Tan +21

Most common SNPs are popularly assumed to be neutral. We here developed novel methods to examine in animal models and humans whether extreme amount of minor alleles (MAs) carried b…

cs.CV2023

Fearless Luminance Adaptation: A Macro-Micro-Hierarchical Transformer for Exposure Correction

Gehui Li, Jinyuan Liu, Long Ma +3

Photographs taken with less-than-ideal exposure settings often display poor visual quality. Since the correction procedures vary significantly, it is difficult for a single neural…

cs.CV2023

From Text to Pixels: A Context-Aware Semantic Synergy Solution for Infrared and Visible Image Fusion

Xingyuan Li, Yang Zou, Jinyuan Liu +4

With the rapid progression of deep learning technologies, multi-modality image fusion has become increasingly prevalent in object detection tasks. Despite its popularity, the inher…

cs.CV2023

MonoTDP: Twin Depth Perception for Monocular 3D Object Detection in Adverse Scenes

Xingyuan Li, Jinyuan Liu, Yixin Lei +3

3D object detection plays a crucial role in numerous intelligent vision systems. Detection in the open world inevitably encounters various adverse scenes, such as dense fog, heavy…

cs.CV2026

Bilevel Layer-Positioning LoRA for Real Image Dehazing

Yan Zhang, Long Ma, Yuxin Feng +3

Learning-based real image dehazing methods have achieved notable progress, yet they still face adaptation challenges in diverse real haze scenes. These challenges mainly stem from…

cond-mat.supr-con2011

Local spin fluctuations in iron-based superconductors: 77Se and 87Rb NMR measurements of Tl0.47Rb0.34Fe1.63Se2

Long Ma, G. F. Ji, J. Dai +5

We report nuclear magnetic resonance (NMR) studies of the intercalated iron selenide superconductor (Tl, Rb)FeSe ( K). Single-crystal measurements up to…

cs.CV2026

Bridging Human Evaluation to Infrared and Visible Image Fusion

Jinyuan Liu, Xingyuan Li, Qingyun Mei +5

Infrared and visible image fusion (IVIF) integrates complementary modalities to enhance scene perception. Current methods predominantly focus on optimizing handcrafted losses and o…

physics.chem-ph2021

High-Performance Nanofluidic Osmotic Power Generation Enabled by Exterior Surface Charges under the Natural Salt Gradient

Long Ma, Zhongwu Li, Zhishan Yuan +3

High-performance osmotic energy conversion (OEC) requires both high ionic selectivity and permeability in nanopores. Here, through systematical explorations of influences from indi…

cs.SD2026

STEB: A Speech-to-Speech Translation Expressiveness Benchmark for Evaluating Beyond Translation Fidelity

Sitong Cheng, Weizhen Bian, Songjun Cao +9

Speech-to-speech translation (S2ST) should preserve not only lexical meaning, but also expressive attributes: emotion, scenario style (e.g., news reporting vs. dramatic dialogue),…

cs.CV2026

Beyond Face Swapping: A Diffusion-Based Digital Human Benchmark for Multimodal Deepfake Detection

Jiaxin Liu, Jia Wang, Saihui Hou +5

In recent years, the explosive advancement of deepfake technology has posed a critical and escalating threat to public security: diffusion-based digital human generation. Unlike tr…

cs.CV2026

NTIRE 2026 3D Restoration and Reconstruction in Real-world Adverse Conditions: RealX3D Challenge Results

Shuhong Liu, Chenyu Bao, Ziteng Cui +103

This paper presents a comprehensive review of the NTIRE 2026 3D Restoration and Reconstruction (3DRR) Challenge, detailing the proposed methods and results. The challenge seeks to…

physics.chem-ph2025

Ionic current rectification under concentration gradients and its application in evaluating surface charge properties of micropores

Long Ma, Hongwen Zhang, Bowen Ai +3

Ionic current rectification (ICR) induced by electroosmotic flow (EOF) under concentration gradients can find many applications in micro/nanofluidic sensing and ionic circuits. Her…

gr-qc2023

Transient analysis of arm locking controller

Yi Zhang, Mingzhe Li, Tong Wang +4

Arm locking is one of the key technologies to suppress the laser phase noise in spaced-based gravitational waves observatories. Since arm locking was proposed, phase margin criteri…

nucl-th2021

Parton collisional effect on the conversion of geometry eccentricities into momentum anisotropies in relativistic heavy-ion collisions

Long Ma, Guo-Liang Ma, Yu-Gang Ma

We explore parton collisional effects on the conversion of geometry eccentricities into azimuthal anisotropies in Pb+Pb collisions at = 5.02 TeV using a multi-phase…

physics.chem-ph2023

Theoretical prediction of diffusive ionic current through nanopores under salt gradients

Long Ma, Zihao Gao, Jia Man +3

In charged nanopores, ionic diffusion current reflects the ionic selectivity and ionic permeability of nanopores which determines the performance of osmotic energy conversion, i.e.…

cs.SD2023

Two Stage Contextual Word Filtering for Context bias in Unified Streaming and Non-streaming Transducer

Zhanheng Yang, Sining Sun, Xiong Wang +3

It is difficult for an E2E ASR system to recognize words such as entities appearing infrequently in the training data. A widely used method to mitigate this issue is feeding contex…

q-bio.GN2013

An Efficient Sufficient Dimension Reduction Method for Identifying Genetic Variants of Clinical Significance

Momiao Xiong, Long Ma

Fast and cheaper next generation sequencing technologies will generate unprecedentedly massive and highly-dimensional genomic and epigenomic variation data. In the near future, a r…

physics.soc-ph2021

Two-population SIR model and strategies to reduce mortality in pandemics

Long Ma, Maksim Kitsak, Piet Van Mieghem

Despite many studies on the transmission mechanism of the Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), it remains still challenging to efficiently reduce mortality…

physics.chem-ph2024

Dynamic Response of Ionic Current in Conical Nanopores

Zhe Liu, Long Ma, Hongwen Zhang +4

Ionic current rectification (ICR) of charged conical nanopores has various applications in fields including nanofluidics, bio-sensing, and energy conversion, whose function is clos…

cond-mat.supr-con2013

Review of nuclear magnetic resonance studies on iron-based superconductors

Long Ma, Weiqiang Yu

The newly discovered iron-based superconductors have triggered renewed enormous research interest in the condensed matter physics community. Nuclear magnetic resonance (NMR) is a l…

cond-mat.mes-hall2025

Radiative Heat Transfer and 2D Transition Metal Dichalcogenide Materials

Long Ma, Dai-Nam Le, Lilia M. Woods

Radiative heat transfer is of great interest from a fundamental point of view and for energy harvesting applications. This is a material dependent phenomenon where confined plasmon…

cs.CV2024

Seeing Text in the Dark: Algorithm and Benchmark

Chengpei Xu, Hao Fu, Long Ma +6

Localizing text in low-light environments is challenging due to visual degradations. Although a straightforward solution involves a two-stage pipeline with low-light image enhancem…

physics.chem-ph2024

Modulation of ionic current rectification in short bipolar nanopores

Hongwen Zhang, Long Ma, Chao Zhang +1

Bipolar nanopores, with asymmetric charge distributions, can induce significant ionic current rectification (ICR) at ultra-short lengths, finding potential applications in nanoflui…

physics.soc-ph2015

The spreading ability of nodes towards localized targets in complex networks

Ye Sun, Long Ma, An Zeng +1

As an important type of dynamics on complex networks, spreading is widely used to model many real processes such as the epidemic contagion and information propagation. One of the m…

cs.SD2025

FreeCodec: A disentangled neural speech codec with fewer tokens

Youqiang Zheng, Weiping Tu, Yueteng Kang +5

Neural speech codecs have gained great attention for their outstanding reconstruction with discrete token representations. It is a crucial component in generative tasks such as spe…

cs.CV2025

DCEvo: Discriminative Cross-Dimensional Evolutionary Learning for Infrared and Visible Image Fusion

Jinyuan Liu, Bowei Zhang, Qingyun Mei +6

Infrared and visible image fusion integrates information from distinct spectral bands to enhance image quality by leveraging the strengths and mitigating the limitations of each mo…

physics.ins-det2015

A new method of energy calibration of position-sensitive silicon detector

Mingdao Sun, Tianheng Huang, Zhong Liu +9

An improved method of energy calibration of position-sensitive silicon detector is presented. Instead of the parabolic function used in traditional method, a new function describin…

cond-mat.str-el2020

NMR study of the spin correlations in the armchair chain NiNbBO

K. Y. Zeng, Long Ma, L. M. Xu +3

We report our nuclear magnetic resonance (NMR) study on the structurally spin chain compound NiNbBO with complex magnetic coupling. The antiferromagnetic transition is moni…

cs.CV2025

DifIISR: A Diffusion Model with Gradient Guidance for Infrared Image Super-Resolution

Xingyuan Li, Zirui Wang, Yang Zou +5

Infrared imaging is essential for autonomous driving and robotic operations as a supportive modality due to its reliable performance in challenging environments. Despite its popula…

cs.CV2025

CoA: Towards Real Image Dehazing via Compression-and-Adaptation

Long Ma, Yuxin Feng, Yan Zhang +5

Learning-based image dehazing algorithms have shown remarkable success in synthetic domains. However, real image dehazing is still in suspense due to computational resource constra…

quant-ph2017

Higher Order Mode Entanglement in a Type II Optical Parametric Oscillator

Jun Guo, Chunxiao Cai, Long Ma +3

Nonclassical beams in high order spatial modes have attracted much interest but they exhibit much less squeezing and entanglement than the fundamental spatial modes, limiting their…

cs.CV2025

PCDreamer: Point Cloud Completion Through Multi-view Diffusion Priors

Guangshun Wei, Yuan Feng, Long Ma +3

This paper presents PCDreamer, a novel method for point cloud completion. Traditional methods typically extract features from partial point clouds to predict missing regions, but t…

cs.CL2020

Multi-head Monotonic Chunkwise Attention For Online Speech Recognition

Baiji Liu, Songjun Cao, Sining Sun +2

The attention mechanism of the Listen, Attend and Spell (LAS) model requires the whole input sequence to calculate the attention context and thus is not suitable for online speech…

cs.CV2022

Toward Fast, Flexible, and Robust Low-Light Image Enhancement

Long Ma, Tengyu Ma, Risheng Liu +2

Existing low-light image enhancement techniques are mostly not only difficult to deal with both visual quality and computational efficiency but also commonly invalid in unknown com…

cs.CV2021

Task-Oriented Convex Bilevel Optimization with Latent Feasibility

Risheng Liu, Long Ma, Xiaoming Yuan +2

This paper firstly proposes a convex bilevel optimization paradigm to formulate and optimize popular learning and vision problems in real-world scenarios. Different from convention…

cs.CY2025

Social World Model-Augmented Mechanism Design Policy Learning

Xiaoyuan Zhang, Yizhe Huang, Chengdong Ma +6

Designing adaptive mechanisms to align individual and collective interests remains a central challenge in artificial social intelligence. Existing methods often struggle with model…

cs.CV2023

Bilevel Generative Learning for Low-Light Vision

Yingchi Liu, Zhu Liu, Long Ma +4

Recently, there has been a growing interest in constructing deep learning schemes for Low-Light Vision (LLV). Existing techniques primarily focus on designing task-specific and dat…

cs.CV2026

Practical exposure correction via compensation

Long Ma, Nan An, Jinyuan Liu +4

In computer vision, correcting the exposure level is a fundamental task for enhancing the visual quality of observations with inappropriate lightness. However, existing methodologi…

cs.CV2025

Striving for Faster and Better: A One-Layer Architecture with Auto Re-parameterization for Low-Light Image Enhancement

Nan An, Long Ma, Guangchao Han +2

Deep learning-based low-light image enhancers have made significant progress in recent years, with a trend towards achieving satisfactory visual quality while gradually reducing th…

cs.AI2026

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning

Wanshi Xu, Haokun Zhao, Haidong Yuan +2

Chain-of-Thought (CoT) reasoning has extended from purely linguistic domains to multimodal scenarios; however, existing approaches often treat visual inputs as homogeneous or auxil…

physics.chem-ph2024

Ion Transport through Short Nanopores Modulated by Charged Exterior Surfaces

Long Ma, Zhe Liu, Bowen Ai +4

Short nanopores find extensive applications capitalizing on their high throughput and detection resolution. Ionic behaviors through long nanopores are mainly determined by charged…

eess.AS2021

Tiny Transducer: A Highly-efficient Speech Recognition Model on Edge Devices

Yuekai Zhang, Sining Sun, Long Ma

This paper proposes an extremely lightweight phone-based transducer model with a tiny decoding graph on edge devices. First, a phone synchronous decoding (PSD) algorithm based on b…

eess.AS2022

A practical framework for multi-domain speech recognition and an instance sampling method to neural language modeling

Yike Zhang, Xiaobing Feng, Yi Liu +2

Automatic speech recognition (ASR) systems used on smart phones or vehicles are usually required to process speech queries from very different domains. In such situations, a vanill…

cs.LG2026

Thought Purity: A Defense Framework For Chain-of-Thought Attack

Zihao Xue, Zhen Bi, Long Ma +7

Large Reasoning Models (LRMs) leverage Chain-of-Thought (CoT) reasoning to solve complex tasks, but this explicit reasoning process introduces a critical vulnerability: adversarial…

physics.chem-ph2024

Influences of Divalent Ions in Natural Seawater/River Water on Nanofluidic Osmotic Energy Generation

Fenhong Song, Xuan An, Long Ma +2

Besides the dominant NaCl, natural seawater/river water contains trace multivalent ions, which can provide effective screening to surface charges. Here, in both negatively and posi…

cs.CV2018

Learning Converged Propagations with Deep Prior Ensemble for Image Enhancement

Risheng Liu, Long Ma, Yiyang Wang +1

Enhancing visual qualities of images plays very important roles in various vision and learning applications. In the past few years, both knowledge-driven maximum a posterior (MAP)…

quant-ph2021

Markov chains and hitting times for error accumulation in quantum circuits

Long Ma, Jaron Sanders

We study a classical model for the accumulation of errors in multi-qubit quantum computations. By modeling the error process in a quantum computation using two coupled Markov chain…

cs.SD2026

Detect All-Type Deepfake Audio: Wavelet Prompt Tuning for Enhanced Auditory Perception

Yuankun Xie, Ruibo Fu, Zhiyong Wang +5

The rapid advancement of audio generation technologies has escalated the risks of malicious deepfake audio across speech, sound, singing voice, and music, threatening multimedia se…

cond-mat.str-el2021

Local evidence for collective spin excitations in the distorted kagome antiferromagnet PrBWO

K. Y. Zeng, F. Y. Song, Z. M. Tian +8

We report the local probe investigation of a frustrated antiferromagnet PrBWO with the distorted kagome lattice. Absence of magnetic order or spin freezing is indicated by…

cs.SD2026

Leveraging large multimodal models for audio-video deepfake detection: a pilot study

Songjun Cao, Yuqi Li, Yunpeng Luo +2

Audio-visual deepfake detection (AVD) is increasingly important as modern generators can fabricate convincing speech and video. Most current multimodal detectors are small, task-sp…

cs.CV2024

3D-GOI: 3D GAN Omni-Inversion for Multifaceted and Multi-object Editing

Haoran Li, Long Ma, Haolin Shi +4

The current GAN inversion methods typically can only edit the appearance and shape of a single object and background while overlooking spatial information. In this work, we propose…

cs.CV2026

TopoMamba: Topology-Aware Scanning and Fusion for Segmenting Heterogeneous Medical Visual Media

Fuchen Zheng, Chengpei Xu, Long Ma +10

Visual state-space models (SSMs) have shown strong potential for medical image segmentation, yet their effectiveness is often limited by two practical issues: axis-biased scan orde…

cs.AI2026

Thinking with Constructions: A Benchmark and Policy Optimization for Visual-Text Interleaved Geometric Reasoning

Haokun Zhao, Wanshi Xu, Haidong Yuan +3

Geometric reasoning inherently requires "thinking with constructions" -- the dynamic manipulation of visual aids to bridge the gap between problem conditions and solutions. However…

physics.chem-ph2025

Characteristics of mono-, di-, and trivalent cations in electric double layers: a molecular dynamic investigation

Bowen Ai, Zekun Gong, Long Ma +3

Ionic behaviors, including ion distributions and hydration characteristics at solid-liquid interfaces, are important research interests in many important applications, such as elec…

cs.CV2023

Trash to Treasure: Low-Light Object Detection via Decomposition-and-Aggregation

Xiaohan Cui, Long Ma, Tengyu Ma +3

Object detection in low-light scenarios has attracted much attention in the past few years. A mainstream and representative scheme introduces enhancers as the pre-processing for re…

eess.AS2023

DCCRN-KWS: an audio bias based model for noise robust small-footprint keyword spotting

Shubo Lv, Xiong Wang, Sining Sun +2

Real-world complex acoustic environments especially the ones with a low signal-to-noise ratio (SNR) will bring tremendous challenges to a keyword spotting (KWS) system. Inspired by…

cs.CV2023

Multi-interactive Feature Learning and a Full-time Multi-modality Benchmark for Image Fusion and Segmentation

Jinyuan Liu, Zhu Liu, Guanyao Wu +5

Multi-modality image fusion and segmentation play a vital role in autonomous driving and robotic operation. Early efforts focus on boosting the performance for only one task, \emph…

cs.CV2025

From Specificity to Generality: Revisiting Generalizable Artifacts in Detecting Face Deepfakes

Long Ma, Zhiyuan Yan, Jin Xu +5

Detecting deepfakes has been an increasingly important topic, especially given the rapid development of AI generation techniques. In this paper, we ask: How can we build a universa…

cs.SD2022

Conversational Speech Recognition By Learning Conversation-level Characteristics

Kun Wei, Yike Zhang, Sining Sun +2

Conversational automatic speech recognition (ASR) is a task to recognize conversational speech including multiple speakers. Unlike sentence-level ASR, conversational ASR can natura…

cond-mat.supr-con2014

Phase Separation, Competition, and Volume Fraction Control in NaFeCoAs

Long Ma, J. Dai, P. S. Wang +9

We report a detailed nuclear magnetic resonance (NMR) study by combined Na and As measurements over a broad range of doping to map the phase diagram of NaFeCo…

cond-mat.str-el2022

Incommensurate magnetic order in SmBWO with the distorted kagome lattice

K. Y. Zeng, F. Y. Song, L. S. Ling +5

We investigate the magnetic ground state of SmBWO with the distorted kagome lattice. A magnetic phase transition is identified at K from the temperature dependen…

physics.chem-ph2024

Detection of Nanopores with the Scanning Ion Conductance Microscopy: A Simulation Study

Yinghua Qiu, Long Ma, Zhe Liu +3

During the dielectric breakdown process of thin solid-state nanopores, the application of high voltages may cause the formation of multi-nanopores on one chip, which number and siz…

eess.AS2025

PseudoVC: Improving One-shot Voice Conversion with Pseudo Paired Data

Songjun Cao, Qinghua Wu, Jie Chen +2

As parallel training data is scarce for one-shot voice conversion (VC) tasks, waveform reconstruction is typically performed by various VC systems. A typical one-shot VC system com…

cs.CV2020

Retinex-inspired Unrolling with Cooperative Prior Architecture Search for Low-light Image Enhancement

Risheng Liu, Long Ma, Jiaao Zhang +2

Low-light image enhancement plays very important roles in low-level vision field. Recent works have built a large variety of deep learning models to address this task. However, the…

cs.AI2026

Steering LLMs via Scalable Interactive Oversight

Enyu Zhou, Zhiheng Xi, Long Ma +9

As Large Language Models increasingly automate complex, long-horizon tasks such as \emph{vibe coding}, a supervision gap has emerged. While models excel at execution, users often s…

cs.CV2023

Improving Misaligned Multi-modality Image Fusion with One-stage Progressive Dense Registration

Di Wang, Jinyuan Liu, Long Ma +2

Misalignments between multi-modality images pose challenges in image fusion, manifesting as structural distortions and edge ghosts. Existing efforts commonly resort to registering…

cs.CV2023

Bi-level Dynamic Learning for Jointly Multi-modality Image Fusion and Beyond

Zhu Liu, Jinyuan Liu, Guanyao Wu +3

Recently, multi-modality scene perception tasks, e.g., image fusion and scene understanding, have attracted widespread attention for intelligent vision systems. However, early effo…

hep-ph2021

Study of a background reconstruction method for the measurement of D-meson azimuthal angular correlations

Long Ma, Xin Dong, Huan-Zhong Huang +1

We study experimental background reconstruction methods for the measurement of correlation using a PYTHIA simulation. Like-sign and side-band background methods th…

cs.SD2025

MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt

Zhichao Wu, Yueteng Kang, Songjun Cao +3

Most existing Zero-Shot Text-To-Speech(ZS-TTS) systems generate the unseen speech based on single prompt, such as reference speech or text descriptions, which limits their flexibil…

cs.CV2024

HUPE: Heuristic Underwater Perceptual Enhancement with Semantic Collaborative Learning

Zengxi Zhang, Zhiying Jiang, Long Ma +3

Underwater images are often affected by light refraction and absorption, reducing visibility and interfering with subsequent applications. Existing underwater image enhancement met…

cond-mat.supr-con2012

NMR study of superconductivity and spin fluctuations in hole-doped superconductor Ca1-xNaxFe2As2 (Tc =32 K)

Long Ma, J. S. Zhang, D. M. Wang +4

We report both 23Na and 75As NMR studies on hole-doped Ca1-xNaxFe2As2 superconducting single crystals (x\approx 0.67) with Tc =32 K. Singlet superconductivity is suggested by a sha…

physics.chem-ph2024

Influences of Electroosmotic Flow on Ionic Current through Nanopores: a Comprehensive Understanding

Yinghua Qiu, Long Ma

Continuum simulations become an important tool to uncover the mysteries in nanofluidic experiments. However, fluid flow in simulation models is usually unconsidered. Here, systemat…

cond-mat.str-el2024

Bose-Einstein condensation of a two-magnon bound state in a spin-one triangular lattice

Jieming Sheng, Jia-Wei Mei, Le Wang +31

In ordered magnets, the elementary excitations are spin waves (magnons), which obey Bose-Einstein statistics. Similarly to Cooper pairs in superconductors, magnons can be paired in…

physics.soc-ph2020

Network-Based Prediction of the 2019-nCoV Epidemic Outbreak in the Chinese Province Hubei

Bastian Prasse, Massimo A. Achterberg, Long Ma +1

At the moment of writing (12 February, 2020), the future evolution of the 2019-nCoV virus is unclear. Predictions of the further course of the epidemic are decisive to deploy targe…

cs.CV2023

PAIF: Perception-Aware Infrared-Visible Image Fusion for Attack-Tolerant Semantic Segmentation

Zhu Liu, Jinyuan Liu, Benzhuang Zhang +3

Infrared and visible image fusion is a powerful technique that combines complementary information from different modalities for downstream semantic perception tasks. Existing learn…

nucl-ex2017

Measurement of -meson triggered correlations in p+p collisions at RHIC

Long Ma

We report the preliminary results of the azimuthal correlations between mesons and charged hadrons (-h) measured by the STAR experiment in proton+proton collision…

cs.SD2026

Diffusion Reconstruction towards Generalizable Audio Deepfake Detection

Bo Cheng, Songjun Cao, Xiaoming Zhang +3

Achieving robust generalization against unseen attacks remains a challenge in Audio Deepfake Detection (ADD), driven by the rapid evolution of generative models. To address this, w…

quant-ph2020

Generation of the Squeezed State with an Arbitrary Complex Amplitude Distribution

Long Ma, Hui Guo, Hengxin Sun +3

The squeezed state is important in quantum metrology and quantum information. The most effective generation tool known is the optical parametric oscillator (OPO). Currently, only t…

cs.CV2025

Detecting AI-Generated Video via Frame Consistency

Long Ma, Zhiyuan Yan, Qinglang Guo +3

The escalating quality of video generated by advanced video generation methods results in new security challenges, while there have been few relevant research efforts: 1) There is…