papers

Publications (25)

stat.ML2025

Variance Reduction via Resampling and Experience Replay

Jiale Han, Xiaowu Dai, Yuhua Zhu

Experience replay is a foundational technique in reinforcement learning that enhances learning stability by storing past experiences in a replay buffer and reusing them during trai…

cs.CR2026

Beyond A Fixed Seal: Adaptive Stealing Watermark in Large Language Models

Shuhao Zhang, Yuli Chen, Jiale Han +2

Watermarking provides a critical safeguard for large language model (LLM) services by facilitating the detection of LLM-generated text. Correspondingly, stealing watermark algorith…

cs.AI2026

SimDiff: Depth Pruning via Similarity and Difference

Yuli Chen, Shuhao Zhang, Fanshen Meng +4

Depth pruning improves the deployment efficiency of large language models (LLMs) by identifying and removing redundant layers. A widely accepted standard for this identification pr…

cs.LG2026

Auto-bidding under Return-on-Spend Constraints with Uncertainty Quantification

Jiale Han, Chun Gan, Chengcheng Zhang +4

Auto-bidding systems are widely used in advertising to automatically determine bid values under constraints such as total budget and Return-on-Spend (RoS) targets. Existing works o…

cs.CE2025

Aethorix v1.0: An Integrated Scientific AI Agent for Scalable Inorganic Materials Innovation and Industrial Implementation

Yingjie Shi, Yiru Gong, Yiqun Su +3

Artificial Intelligence (AI) is redefining the frontiers of scientific domains, ranging from drug discovery to meteorological modeling, yet its integration within industrial manufa…

cs.LG2025

Incentivizing Truthful Language Models via Peer Elicitation Games

Baiting Chen, Tong Zhu, Jiale Han +3

Large Language Models (LLMs) have demonstrated strong generative capabilities but remain prone to inconsistencies and hallucinations. We introduce Peer Elicitation Games (PEG), a t…

cs.GT2025

Online Auction Design Using Distribution-Free Uncertainty Quantification with Applications to E-Commerce

Jiale Han, Xiaowu Dai

Online auction is a cornerstone of e-commerce, and a key challenge is designing incentive-compatible mechanisms that maximize expected revenue. Existing approaches often assume kno…

cs.RO2025

A Multi-view Landmark Representation Approach with Application to GNSS-Visual-Inertial Odometry

Tong Hua, Jiale Han, Wei Ouyang

Invariant Extended Kalman Filter (IEKF) has been a significant technique in vision-aided sensor fusion. However, it usually suffers from high computational burden when jointly opti…

cs.CL2021

Integrating Subgraph-aware Relation and DirectionReasoning for Question Answering

Xu Wang, Shuai Zhao, Bo Cheng +5

Question Answering (QA) models over Knowledge Bases (KBs) are capable of providing more precise answers by utilizing relation information among entities. Although effective, most o…

cs.CL2025

DialogueAgents: A Hybrid Agent-Based Speech Synthesis Framework for Multi-Party Dialogue

Xiang Li, Duyi Pan, Hongru Xiao +5

Speech synthesis is crucial for human-computer interaction, enabling natural and intuitive communication. However, existing datasets involve high construction costs due to manual a…

cs.GT2026

A Robust Multi-Item Auction Design with Statistical Learning

Jiale Han, Xiaowu Dai

We propose a novel statistical learning method for multi-item auctions that incorporates credible intervals. Our approach employs nonparametric density estimation to estimate credi…

cs.CL2022

Generative Prompt Tuning for Relation Classification

Jiale Han, Shuai Zhao, Bo Cheng +2

Using prompts to explore the knowledge contained within pre-trained language models for downstream tasks has now become an active topic. Current prompt tuning methods mostly conver…

cs.GT2026

Mechanism Design for Quality-Preserving LLM Advertising

Jiale Han, Xiaowu Dai

Embedding advertisements into large language model (LLM) outputs introduces a fundamental tension: revenue optimization can distort content and degrade user experience. Existing ap…

cs.IR2025

Adapting General-Purpose Embedding Models to Private Datasets Using Keyword-based Retrieval

Yubai Wei, Jiale Han, Yi Yang

Text embedding models play a cornerstone role in AI applications, such as retrieval-augmented generation (RAG). While general-purpose text embedding models demonstrate strong perfo…

cs.SD2024

CM-TTS: Enhancing Real Time Text-to-Speech Synthesis Efficiency through Weighted Samplers and Consistency Models

Xiang Li, Fan Bu, Ambuj Mehrish +4

Neural Text-to-Speech (TTS) systems find broad applications in voice assistants, e-learning, and audiobook creation. The pursuit of modern models, like Diffusion Models (DMs), hold…

cs.RO2024

An Immediate Update Strategy of Multi-State Constraint Kalman Filter

Qingchao Zhang, Wei Ouyang, Jiale Han +3

The lightweight Multi-state Constraint Kalman Filter (MSCKF) has been well-known for its high efficiency, in which the delayed update has been usually adopted since its proposal. T…

cs.CL2026

HiMed: Incentivizing Hindi Reasoning in Medical LLMs

Dingfeng Jiang, Han Yan, Chenze Ma +12

Medical large language models hold promise for reducing healthcare disparities, yet Hindi remains severely underrepresented. While medical LLMs excel in high-resource languages, th…

eess.SY2025

CT-ESKF: A General Framework of Covariance Transformation-Based Error-State Kalman Filter

Jiale Han, Wei Ouyang, Maoran Zhu +1

Invariant extended Kalman filter (InEKF) possesses excellent trajectory-independent property and better consistency compared to conventional extended Kalman filter (EKF). However,…

cs.CL2025

DLP: Dynamic Layerwise Pruning in Large Language Models

Yuli Chen, Bo Cheng, Jiale Han +3

Pruning has recently been widely adopted to reduce the parameter scale and improve the inference efficiency of Large Language Models (LLMs). Mainstream pruning techniques often rel…

cs.CL2021

Exploring Task Difficulty for Few-Shot Relation Extraction

Jiale Han, Bo Cheng, Wei Lu

Few-shot relation extraction (FSRE) focuses on recognizing novel relations by learning with merely a handful of annotated instances. Meta-learning has been widely adopted for such…

cs.AI2026

Human or Machine? A Preliminary Turing Test for Speech-to-Speech Interaction

Xiang Li, Jiabao Gao, Sipei Lin +5

The pursuit of human-like conversational agents has long been guided by the Turing test. For modern speech-to-speech (S2S) systems, a critical yet unanswered question is whether th…

cs.IR2025

RAG Meets Temporal Graphs: Time-Sensitive Modeling and Retrieval for Evolving Knowledge

Jiale Han, Austin Cheung, Yubai Wei +4

Knowledge is inherently time-sensitive and continuously evolves over time. Although current Retrieval-Augmented Generation (RAG) systems enrich LLMs with external knowledge, they l…

cs.CR2025

CEFW: A Comprehensive Evaluation Framework for Watermark in Large Language Models

Shuhao Zhang, Bo Cheng, Jiale Han +4

Text watermarking provides an effective solution for identifying synthetic text generated by large language models. However, existing techniques often focus on satisfying specific…

cs.CL2024

Making Pre-trained Language Models Better Continual Few-Shot Relation Extractors

Shengkun Ma, Jiale Han, Yi Liang +1

Continual Few-shot Relation Extraction (CFRE) is a practical problem that requires the model to continuously learn novel relations while avoiding forgetting old ones with few label…

cs.CV2025

HoneyImage: Verifiable, Harmless, and Stealthy Dataset Ownership Verification for Image Models

Zhihao Zhu, Jiale Han, Yi Yang

Image-based AI models are increasingly deployed across a wide range of domains, including healthcare, security, and consumer applications. However, many image datasets carry sensit…