papers

Publications (97)

astro-ph.IM2025

Hierarchical search method for gravitational waves from stellar-mass binary black holes in noisy space-based detector data

Yao Fu, Yan Wang, Soumya D. Mohanty

Future space-based laser interferometric detectors, such as LISA, will be able to detect gravitational waves (GWs) generated during the inspiral phase of stellar-mass binary black…

hep-ph2022

ResBos2 and the CDF W Mass Measurement

Joshua Isaacson, Yao Fu, C. -P. Yuan

The recent CDF mass measurement of 80,433 9 MeV is the most precise direct measurement. However, this result deviates from the Standard Model predicted mass of 80,359.1 $…

eess.IV2026

RAM-W600: A Multi-Task Wrist Dataset and Benchmark for Rheumatoid Arthritis

Songxiao Yang, Haolin Wang, Yao Fu +5

Rheumatoid arthritis (RA) is a common autoimmune disease that has been the focus of research in computer-aided diagnosis (CAD) and disease monitoring. In clinical settings, convent…

cs.LG2025

MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Leyang Xue, Yao Fu, Zhan Lu +2

This paper presents MoE-Infinity, an efficient MoE inference system designed for personal machines with limited GPU memory capacity. The key idea for MoE-Infinity is that on person…

hep-ph2026

A Fast Method for Correlated Updates of Proton PDFs and the Strong Coupling

Yao Fu, Carl Schmidt, C. --P. Yuan

We present an extended version of the \texttt{ePump} framework that enables the simultaneous profiling of proton parton distribution functions (PDFs) and the strong coupling

cs.LG2024

Toward Inference-optimal Mixture-of-Expert Large Language Models

Longfei Yun, Yonghao Zhuang, Yao Fu +2

Mixture-of-Expert (MoE) based large language models (LLMs), such as the recent Mixtral and DeepSeek-MoE, have shown great promise in scaling model size without suffering from the q…

quant-ph2022

Experimental quantum advantage with quantum coupon collector

Min-Gang Zhou, Xiao-Yu Cao, Yu-Shuo Lu +6

An increasing number of communication and computational schemes with quantum advantages have recently been proposed, which implies that quantum technology has fertile application p…

cs.CL2024

Interactive and Expressive Code-Augmented Planning with Large Language Models

Anthony Z. Liu, Xinhe Wang, Jacob Sansom +5

Large Language Models (LLMs) demonstrate strong abilities in common-sense reasoning and interactive decision-making, but often struggle with complex, long-horizon planning tasks. R…

cs.AI2025

When Truthful Representations Flip Under Deceptive Instructions?

Xianxuan Long, Yao Fu, Runchao Li +4

Large language models (LLMs) tend to follow maliciously crafted instructions to generate deceptive responses, posing safety challenges. How deceptive instructions alter the interna…

eess.SY2024

Security and Privacy of Digital Twins for Advanced Manufacturing: A Survey

Alexander D. Zemskov, Yao Fu, Runchao Li +9

In Industry 4.0, the digital twin is one of the emerging technologies, offering simulation abilities to predict, refine, and interpret conditions and operations, where it is crucia…

quant-ph2022

Breaking the Rate-Loss Bound of Quantum Key Distribution with Asynchronous Two-Photon Interference

Yuan-Mei Xie, Yu-Shuo Lu, Chen-Xun Weng +7

Twin-field quantum key distribution can overcome the secret key capacity of repeaterless quantum key distribution via single-photon interference. However, to compensate for the cha…

cs.LG2026

Sample-Mean Anchored Thompson Sampling for Offline-to-Online Learning with Distribution Shift

Bochao Li, Yao Fu, Wei Chen +1

Offline-to-online learning aims to improve online decision-making by leveraging offline logged data. A central challenge in this setting is the distribution shift between offline a…

quant-ph2023

Phase-Matching Quantum Key Distribution without Intensity Modulation

Shan-Feng Shao, Xiao-Yu Cao, Yuan-Mei Xie +5

Quantum key distribution provides a promising solution for sharing secure keys between two distant parties with unconditional security. Nevertheless, quantum key distribution is st…

quant-ph2016

Security of quantum key distribution with multiphoton components

Hua-Lei Yin, Yao Fu, Yingqiu Mao +1

Most qubit-based quantum key distribution (QKD) protocols extract the secure key merely from single-photon component of the attenuated lasers. However, with the Scarani-Acin-Ribord…

quant-ph2014

Measurement-device-independent quantum key distribution based on Bell's inequality

Hua-Lei Yin, Yao Fu, Yan-Lin Tang +3

We propose two quantum key distribution (QKD) protocols based on Bell's inequality, which can be considered as modified time-reversed E91 protocol. Similar to the measurement-devic…

quant-ph2025

Multi-field quantum conferencing overcomes the network capacity limit

Yuan-Mei Xie, Yu-Shuo Lu, Yao Fu +2

Quantum conferencing enables multiple nodes within a quantum network to share a secure group key for private message broadcasting. The key rate, however, is limited by the repeater…

cs.CR2021

On the Practicality of Differential Privacy in Federated Learning by Tuning Iteration Times

Yao Fu, Yipeng Zhou, Di Wu +3

In spite that Federated Learning (FL) is well known for its privacy protection when training machine learning models among distributed clients collaboratively, recent studies have…

quant-ph2019

Phase self-aligned continuous-variable measurement-device-independent quantum key distribution

Hua-Lei Yin, Wei Zhu, Yao Fu

Continuous-variable measurement-independent-device quantum key distribution (CV-MDI-QKD) can offer high secure key rate at metropolitan distance and remove all side channel loophol…

cs.CL2023

FiLM: Fill-in Language Models for Any-Order Generation

Tianxiao Shen, Hao Peng, Ruoqi Shen +3

Language models have become the backbone of today's AI systems. However, their predominant left-to-right generation limits the use of bidirectional context, which is essential for…

cs.CL2022

Few-shot Subgoal Planning with Language Models

Lajanugen Logeswaran, Yao Fu, Moontae Lee +1

Pre-trained large language models have shown successful progress in many language understanding benchmarks. This work explores the capability of these models to predict actionable…

quant-ph2023

Scalable High-Rate Twin-Field Quantum Key Distribution Networks without Constraint of Probability and Intensity

Yuan-Mei Xie, Chen-Xun Weng, Yu-Shuo Lu +4

Implementation of a twin-field quantum key distribution network faces limitations, including the low tolerance of interference errors for phase-matching type protocols and the stri…

hep-ph2023

Improving ResBos for the precision needs of the LHC

Joshua Isaacson, Yao Fu, C. -P. Yuan

The resummation calculation (ResBos) is a widely used tool for the simulation of single vector boson production at colliders. In this work, we develop a significant improvement ove…

cs.CL2022

Data-to-text Generation with Variational Sequential Planning

Ratish Puduppully, Yao Fu, Mirella Lapata

We consider the task of data-to-text generation, which aims to create textual output from non-linguistic input. We focus on generating long-form text, i.e., documents with multiple…

cs.CV2026

Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos

Shoubin Yu, Lei Shu, Antoine Yang +6

Multimodal AI agents are increasingly automating complex real-world workflows that involve online web execution. However, current web-agent benchmarks suffer from a critical limita…

cs.LG2025

HybridServe: Efficient Serving of Large AI Models with Confidence-Based Cascade Routing

Leyang Xue, Yao Fu, Luo Mai +1

Giant Deep Neural Networks (DNNs), have become indispensable for accurate and robust support of large-scale cloud based AI services. However, serving giant DNNs is prohibitively ex…

hep-ph2023

Probing Parton distribution functions at large x via Drell-Yan Forward-Backward Asymmetry

Yao Fu, Raymond Brock, Daniel Hayden +1

The forward-backward asymmetry of the Drell-Yan process in dilepton decays at high invariant masses can be used to probe the parton distribution functions at large x. The behavior…

cs.CL2023

Improving Language Model Negotiation with Self-Play and In-Context Learning from AI Feedback

Yao Fu, Hao Peng, Tushar Khot +1

We study whether multiple large language models (LLMs) can autonomously improve each other in a negotiation game by playing, reflecting, and criticizing. We are interested in this…

quant-ph2023

Breaking Rate-Distance Limitation of Measurement-Device-Independent Quantum Secret Sharing

Chen-Long Li, Yao Fu, Wen-Bo Liu +5

Currently most progresses on quantum secret sharing suffer from rate-distance bound, and thus the key rates are limited. In addition to the limited key rate, the technical difficul…

hep-ph2022

Boost Asymmetry of the diboson productions in pp collisions

Siqi Yang, Mingzhe Xie, Yao Fu +5

We propose the boost asymmetry of the diboson productions in pp collisions as a new experimental observable, which can provide unique information on the proton structure. The boost…

quant-ph2022

Experimental quantum secure network with digital signatures and encryption

Hua-Lei Yin, Yao Fu, Chen-Long Li +6

Cryptography promises four information security objectives, namely, confidentiality, integrity, authenticity, and non-repudiation, to support trillions of transactions annually in…

hep-ph2021

Reduction of the electroweak correlation in the PDF updating by using the forward-backward asymmetry of Drell-Yan process

Siqi Yang, Yao Fu, Minghui Liu +6

We propose a new observable for the measurement of the forward-backward asymmetry in Drell-Yan lepton production. At hadron colliders, the distribution is sensi…

cs.MS2022

TorchOpt: An Efficient Library for Differentiable Optimization

Jie Ren, Xidong Feng, Bo Liu +4

Recent years have witnessed the booming of various differentiable optimization algorithms. These algorithms exhibit different execution patterns, and their execution needs massive…

quant-ph2026

Repeater-like asynchronous measurement-device-independent quantum conference key agreement

Yu-Shuo Lu, Hua-Lei Yin, Yuan-Mei Xie +2

Quantum conference key agreement enables secure communication among multiple parties by leveraging multipartite entanglement, which is expected to play a crucial role in future qua…

cs.CL2025

FAEDKV: Infinite-Window Fourier Transform for Unbiased KV Cache Compression

Runchao Li, Yao Fu, Mu Sheng +3

The efficacy of Large Language Models (LLMs) in long-context tasks is often hampered by the substantial memory footprint and computational demands of the Key-Value (KV) cache. Curr…

quant-ph2016

Practical Quantum Digital Signature

Hua-Lei Yin, Yao Fu, Zeng-Bing Chen

Guaranteeing nonrepudiation, unforgeability as well as transferability of a signature is one of the most vital safeguards in today's e-commerce era. Based on fundamental laws of qu…

quant-ph2021

Secure Quantum Secret Sharing without Signal Disturbance Monitoring

Hua-Lei Yin, Jie Gu, Yuan-Mei Xie +3

Quantum secret sharing (QSS) is an essential primitive for the future quantum internet, which promises secure multiparty communication. However, developing a large-scale QSS networ…

cs.CL2024

DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Guangxuan Xiao, Jiaming Tang, Jingwei Zuo +5

Deploying long-context large language models (LLMs) is essential but poses significant computational and memory challenges. Caching all Key and Value (KV) states across all attenti…

cs.CL2025

Training Foundation Models on a Full-Stack AMD Platform: Compute, Networking, and System Design

Quentin Anthony, Yury Tokpanov, Skyler Szot +18

We report on the first large-scale mixture-of-experts (MoE) pretraining study on pure AMD hardware, utilizing both MI300X GPUs and Pollara networking. We distill practical guidance…

cs.CL2023

Chain-of-Thought Hub: A Continuous Effort to Measure Large Language Models' Reasoning Performance

Yao Fu, Litu Ou, Mingyu Chen +3

As large language models (LLMs) are continuously being developed, their evaluation becomes increasingly important yet challenging. This work proposes Chain-of-Thought Hub, an open-…

cs.CL2024

Retrieval Head Mechanistically Explains Long-Context Factuality

Wenhao Wu, Yizhong Wang, Guangxuan Xiao +2

Despite the recent progress in long-context language models, it remains elusive how transformer-based models exhibit the capability to retrieve relevant information from arbitrary…

cs.CV2026

RAM-H1200: A Unified Evaluation and Dataset on Hand Radiographs for Rheumatoid Arthritis

Songxiao Yang, Haolin Wang, Yao Fu +9

Rheumatoid arthritis (RA) assessment from hand radiographs requires multi-level analysis and modeling of anatomical structures and fine-grained local pathological changes. However,…

quant-ph2024

Experimental quantum e-commerce

Xiao-Yu Cao, Bing-Hong Li, Yang Wang +3

E-commerce, a type of trading that occurs at a high frequency on the Internet, requires guaranteeing the integrity, authentication and non-repudiation of messages through long dist…

quant-ph2023

Source-independent quantum random number generator against tailored detector blinding attacks

Wen-Bo Liu, Yu-Shuo Lu, Yao Fu +5

Randomness, mainly in the form of random numbers, is the fundamental prerequisite for the security of many cryptographic tasks. Quantum randomness can be extracted even if adversar…

quant-ph2019

Measurement-Device-Independent Twin-Field Quantum Key Distribution

Hua-Lei Yin, Yao Fu

The ultimate aim of quantum key distribution (QKD) is improving the performance of transmission distance and key generation speed. Unfortunately, it is believed to be limited by th…

cs.CL2021

Prototypical Representation Learning for Relation Extraction

Ning Ding, Xiaobin Wang, Yao Fu +7

Recognizing relations between entities is a pivotal task of relational learning. Learning relation representations from distantly-labeled datasets is difficult because of the abund…

cs.CL2020

Paraphrase Generation with Latent Bag of Words

Yao Fu, Yansong Feng, John P. Cunningham

Paraphrase generation is a longstanding important problem in natural language processing. In addition, recent progress in deep generative models has shown promising results on disc…

cs.CL2024

OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models

Fuzhao Xue, Zian Zheng, Yao Fu +4

To help the open-source community have a better understanding of Mixture-of-Experts (MoE) based large language models (LLMs), we train and release OpenMoE, a series of fully open-s…

cs.LG2025

MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems

Yinsicheng Jiang, Yao Fu, Yeqi Huang +13

The sparse Mixture-of-Experts (MoE) architecture is increasingly favored for scaling Large Language Models (LLMs) efficiently, but it depends on heterogeneous compute and memory re…

quant-ph2015

Long-Distance Measurement-Device-Independent Multiparty Quantum Communication

Yao Fu, Hua-Lei Yin, Teng-Yun Chen +1

The Greenberger-Horne-Zeilinger (GHZ) entanglement, originally introduced to uncover the extreme violation of local realism against quantum mechanics, is an important resource for…

quant-ph2023

One-Time Universal Hashing Quantum Digital Signatures without Perfect Keys

Bing-Hong Li, Yuan-Mei Xie, Xiao-Yu Cao +4

Quantum digital signatures (QDS), generating correlated bit strings among three remote parties for signatures through quantum law, can guarantee non-repudiation, authenticity, and…

cs.LG2026

Leveraging Extragradient for Effective Sharpness-Aware Minimization in Deep Learning

Yao Fu, Chunxia Zhang, Junmin Liu +3

Generalization remains a pivotal challenge in deep learning, where traditional optimizers like Stochastic Gradient Descent (SGD) often converge to sharp minima, leading to overfitt…

cs.LG2021

Optimizing the Numbers of Queries and Replies in Federated Learning with Differential Privacy

Yipeng Zhou, Xuezheng Liu, Yao Fu +3

Federated learning (FL) empowers distributed clients to collaboratively train a shared machine learning model through exchanging parameter information. Despite the fact that FL can…

cs.LG2024

ServerlessLLM: Low-Latency Serverless Inference for Large Language Models

Yao Fu, Leyang Xue, Yeqi Huang +4

This paper presents ServerlessLLM, a distributed system designed to support low-latency serverless inference for Large Language Models (LLMs). By harnessing the substantial near-GP…

cs.CL2023

Complexity-Based Prompting for Multi-Step Reasoning

Yao Fu, Hao Peng, Ashish Sabharwal +2

We study the task of prompting large-scale language models to perform multi-step reasoning. Existing work shows that when prompted with a chain of thoughts (CoT), sequences of shor…

cs.CL2024

Long Context Alignment with Short Instructions and Synthesized Positions

Wenhao Wu, Yizhong Wang, Yao Fu +3

Effectively handling instructions with extremely long context remains a challenge for Large Language Models (LLMs), typically necessitating high-quality long data and substantial c…

cs.LG2025

Pruning Weights but Not Truth: Safeguarding Truthfulness While Pruning LLMs

Yao Fu, Runchao Li, Xianxuan Long +4

Neural network pruning has emerged as a promising approach for deploying LLMs in low-resource scenarios while preserving downstream task performance. However, for the first time, w…

cs.CL2021

Probing BERT in Hyperbolic Spaces

Boli Chen, Yao Fu, Guangwei Xu +4

Recently, a variety of probing tasks are proposed to discover linguistic properties learned in contextualized word embeddings. Many of these works implicitly assume these embedding…

quant-ph2022

Neural network-based prediction of the secret-key rate of quantum key distribution

Min-Gang Zhou, Zhi-Ping Liu, Wen-Bo Liu +6

Numerical methods are widely used to calculate the secure key rate of many quantum key distribution protocols in practice, but they consume many computing resources and are too tim…

quant-ph2014

Long distance measurement-device-independent quantum key distribution with coherent-state superpositions

Hua-Lei Yin, Wen-Fei Cao, Yao Fu +4

Measurement-device-independent quantum key distribution (MDI-QKD) with decoy-state method is believed to be securely applied to defeat various hacking attacks in practical quantum…

cs.CL2020

Nested Named Entity Recognition with Partially-Observed TreeCRFs

Yao Fu, Chuanqi Tan, Mosha Chen +2

Named entity recognition (NER) is a well-studied task in natural language processing. However, the widely-used sequence labeling framework is difficult to detect entities with nest…

cs.LG2023

Go Beyond Imagination: Maximizing Episodic Reachability with World Models

Yao Fu, Run Peng, Honglak Lee

Efficient exploration is a challenging topic in reinforcement learning, especially for sparse reward tasks. To deal with the reward sparsity, people commonly apply intrinsic reward…

cs.CL2021

Noisy-Labeled NER with Confidence Estimation

Kun Liu, Yao Fu, Chuanqi Tan +4

Recent studies in deep learning have shown significant progress in named entity recognition (NER). Most existing works assume clean data annotation, yet a fundamental challenge in…

cs.CL2020

Latent Template Induction with Gumbel-CRFs

Yao Fu, Chuanqi Tan, Bin Bi +3

Learning to control the structure of sentences is a challenging problem in text generation. Existing work either relies on simple deterministic approaches or RL-based hard structur…

cs.CL2024

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models

Yao Fu, Yin Yu, Xiaotian Han +4

Knowledge distillation (KD) has become a widely adopted approach for compressing large language models (LLMs) to reduce computational costs and memory footprints. However, the avai…

cs.CL2019

Rethinking Text Attribute Transfer: A Lexical Analysis

Yao Fu, Hao Zhou, Jiaze Chen +1

Text attribute transfer is modifying certain linguistic attributes (e.g. sentiment, style, authorship, etc.) of a sentence and transforming them from one type to another. In this p…

cond-mat.mtrl-sci2025

Room temperature optical control of spin states in organic diradicals

Rituparno Chowdhury, Alistair Inglis, Lucy E. Walker +13

We report a family of luminescent alternant diradicals which, at room temperature, support a ground-state spin-triplet, near-unity photoluminescence quantum yields, and optical spi…

quant-ph2023

Asynchronous measurement-device-independent quantum key distribution with hybrid source

Jun-Lin Bai, Yuan-Mei Xie, Yao Fu +2

The linear constraint of secret key rate capacity is overcome by the tiwn-field quantum key distribution (QKD). However, the complex phase-locking and phase-tracking technique requ…

cs.CL2023

C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models

Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu +10

New NLP benchmarks are urgently needed to align with the rapid development of large language models (LLMs). We present C-Eval, the first comprehensive Chinese evaluation suite desi…

quant-ph2022

All-Photonic Quantum Repeater for Multipartite Entanglement Generation

Chen-Long Li, Yao Fu, Wen-Bo Liu +5

Quantum network applications like distributed quantum computing and quantum secret sharing present a promising future network equipped with quantum resources. Entanglement generati…

quant-ph2024

Source-independent quantum secret sharing with entangled photon pair networks

Yi-Ran Xiao, Zhao-Ying Jia, Yu-Chen Song +4

The large-scale deployment of quantum secret sharing (QSS) in quantum networks is currently challenging due to the requirements for the generation and distribution of multipartite…

cs.CL2023

MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning

Xiang Yue, Xingwei Qu, Ge Zhang +5

We introduce MAmmoTH, a series of open-source large language models (LLMs) specifically tailored for general math problem-solving. The MAmmoTH models are trained on MathInstruct, o…

cs.LG2023

To Repeat or Not To Repeat: Insights from Scaling LLM under Token-Crisis

Fuzhao Xue, Yao Fu, Wangchunshu Zhou +2

Recent research has highlighted the importance of dataset size in scaling language models. However, large language models (LLMs) are notoriously token-hungry during pre-training, a…

cs.AI2025

Quantized but Deceptive? A Multi-Dimensional Truthfulness Evaluation of Quantized LLMs

Yao Fu, Xianxuan Long, Runchao Li +5

Quantization enables efficient deployment of large language models (LLMs) in resource-constrained environments by significantly reducing memory and computation costs. While quantiz…

cs.LG2024

Challenges in Deploying Long-Context Transformers: A Theoretical Peak Performance Analysis

Yao Fu

Transformer-based long context generative models power emerging AI applications like hour-long video understanding and project-level coding agent. Deploying long context transforme…

cs.CL2024

AutoGuide: Automated Generation and Selection of Context-Aware Guidelines for Large Language Model Agents

Yao Fu, Dong-Ki Kim, Jaekyeom Kim +4

Recent advances in large language models (LLMs) have empowered AI agents capable of performing various sequential decision-making tasks. However, effectively guiding LLMs to perfor…

cs.LG2025

MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems

Yinsicheng Jiang, Yao Fu, Yeqi Huang +13

The sparse Mixture-of-Experts (MoE) architecture is increasingly favored for scaling Large Language Models (LLMs) efficiently, but it depends on heterogeneous compute and memory re…

cs.LG2022

Scaling Structured Inference with Randomization

Yao Fu, John P. Cunningham, Mirella Lapata

Deep discrete structured models have seen considerable progress recently, but traditional inference using dynamic programming (DP) typically works with a small number of states (le…

hep-ex2022

Measurement of the proton structure parameters in the forward-backward charge asymmetry

Mingzhe Xie, Siqi Yang, Yao Fu +5

The forward-backward asymmetry () in the Drell-Yan process $pp/p\bar p \to Z/γ^* \to \ell^+\ell^-$ is sensitive to the proton structure information. Such information has b…

hep-ph2026

Further Reduction of the PDF Uncertainty in the High-Mass Drell-Yan Spectrum Utilizing Neutral and Charged Current Inputs

Yao Fu, Raymond Brock, Daniel Hayden +1

Uncertainties in the parametrization of Parton Distribution Functions are a serious limiting systematic uncertainty in Large Hadron Collider searches for Beyond the Standard Model…

cs.CL2022

Latent Topology Induction for Understanding Contextualized Representations

Yao Fu, Mirella Lapata

In this work, we study the representation space of contextualized embeddings and gain insight into the hidden topology of large language models. We show there exists a network of l…

cs.CL2024

Critical Data Size of Language Models from a Grokking Perspective

Xuekai Zhu, Yao Fu, Bowen Zhou +1

We explore the critical data size in language models, a threshold that marks a fundamental shift from quick memorization to slow generalization. We formalize the phase transition u…

hep-ex2020

Reduction of PDF uncertainty in the measurement of the weak mixing angle at the ATLAS experiment

Yao Fu, Siqi Yang, Minghui Liu +5

We investigate the parton distribution function (PDF) uncertainty in the measurement of the effective weak mixing angle at the CERN Large Hadron Coll…

hep-ph2022

Factorization of the forward-backward charge asymmetry and measurements of the weak mixing angle and proton structure at hadron colliders

Siqi Yang, Yao Fu, Minghui Liu +3

The forward-backward charge asymmetry (AFB) at hadron colliders is sensitive to both the electroweak (EW) symmetry breaking represented by the effective weak mixing angle, and the…

math.ST2025

Zero-Order Sharpness-Aware Minimization

Yao Fu, Yihang Jin, Chunxia Zhang +3

Prompt learning has become a key method for adapting large language models to specific tasks with limited data. However, traditional gradient-based optimization methods for tuning…

quant-ph2022

Experimental measurement-device-independent type quantum key distribution with flawed and correlated sources

Jie Gu, Xiao-Yu Cao, Yao Fu +4

The security of quantum key distribution (QKD) is severely threatened by discrepancies between realistic devices and theoretical assumptions. Recently, a significant framework call…

cs.DC2026

ProTrain: Efficient LLM Training via Memory-Aware Techniques

Hanmei Yang, Jin Zhou, Yao Fu +4

Memory pressure has emerged as a dominant constraint in scaling the training of large language models (LLMs), particularly in resource-constrained environments. While modern framew…

cond-mat.mtrl-sci2024

Optical read and write of spin states in organic diradicals

Rituparno Chowdhury, Petri Murto, Naitik A. Panjwani +17

Optical control and read-out of the ground state spin structure has been demonstrated for defect states in crystalline semiconductors, including the diamond NV- center, and these a…

cs.CL2023

Decomposed Prompting: A Modular Approach for Solving Complex Tasks

Tushar Khot, Harsh Trivedi, Matthew Finlayson +4

Few-shot prompting is a surprisingly powerful way to use Large Language Models (LLMs) to solve various tasks. However, this approach struggles as the task complexity increases or w…

quant-ph2016

Detector-decoy quantum key distribution without monitoring signal disturbance

Hua-Lei Yin, Yao Fu, Yingqiu Mao +1

The round-robin differential phase-shift quantum key distribution protocol provides a secure way to exchange private information without monitoring conventional disturbances and st…

quant-ph2023

Experimental quantum secret sharing based on phase encoding of coherent states

Ao Shen, Xiao-Yu Cao, Yang Wang +6

Quantum secret sharing (QSS) is one of the basic communication primitives in future quantum networks which addresses part of the basic cryptographic tasks of multiparty communicati…

quant-ph2016

Experimental Quantum Digital Signature over 102 km

Hua-Lei Yin, Yao Fu, Hui Liu +10

Quantum digital signature (QDS) is an approach to guarantee the nonrepudiation, unforgeability and transferability of a signature with the information-theoretical security. All pre…

cs.CL2022

Just-DREAM-about-it: Figurative Language Understanding with DREAM-FLUTE

Yuling Gu, Yao Fu, Valentina Pyatkin +3

Figurative language (e.g., "he flew like the wind") is challenging to understand, as it is hard to tell what implicit information is being conveyed from the surface form alone. We…

cs.CL2024

Data Engineering for Scaling Language Models to 128K Context

Yao Fu, Rameswar Panda, Xinyao Niu +4

We study the continual pretraining recipe for scaling language models' context lengths to 128K, with a focus on data engineering. We hypothesize that long context modeling, in part…

quant-ph2023

Breaking universal limitations on quantum conference key agreement without quantum memory

Chen-Long Li, Yao Fu, Wen-Bo Liu +5

Quantum conference key agreement is an important cryptographic primitive for future quantum network. Realizing this primitive requires high-brightness and robust multiphoton entang…

quant-ph2014

Violations of entropic Bell inequalities with coarse-grained quadrature measurements for continuous-variable states

Zeng-Bing Chen, Yao Fu, Yu-Kang Zhao

It is a long-standing belief, as pointed out by Bell in 1986, that it is impossible to use a two-mode Gaussian state possessing a positive-definite Wigner function to demonstrate n…

cs.CL2023

Specializing Smaller Language Models towards Multi-Step Reasoning

Yao Fu, Hao Peng, Litu Ou +2

The surprising ability of Large Language Models (LLMs) to perform well on complex reasoning with only few-shot chain-of-thought prompts is believed to emerge only in very large-sca…

quant-ph2024

Efficient source-independent quantum conference key agreement

Yu Bao, Yi-Ran Xiao, Yu-Chen Song +4

Quantum conference key agreement (QCKA) enables the unconditional secure distribution of conference keys among multiple participants. Due to challenges in high-fidelity preparation…