papers

Publications (34)

cs.CV2024

DiaLoc: An Iterative Approach to Embodied Dialog Localization

Chao Zhang, Mohan Li, Ignas Budvytis +1

Multimodal learning has advanced the performance for many vision-language tasks. However, most existing works in embodied dialog research focus on navigation and leave the localiza…

eess.AS2021

Head-synchronous Decoding for Transformer-based Streaming ASR

Mohan Li, Catalin Zorila, Rama Doddipatla

Online Transformer-based automatic speech recognition (ASR) systems have been extensively studied due to the increasing demand for streaming applications. Recently proposed Decoder…

econ.GN2024

Economic Forces in Stock Returns

Yue Chen, Mohan Li

When analyzing the components influencing the stock prices, it is commonly believed that economic activities play an important role. More specifically, asset prices are more sensit…

cs.AI2026

FIDES: Faithful Inference via Deep Evidence Signals for Retrieval-Memory Conflict in RAG

Zhe Yu, Wenpeng Xing, Tiancheng Zhao +3

When retrieved evidence contradicts parametric memory, language models frequently ignore context and default to memorized priors -- a failure that undermines the core purpose of re…

cs.LG2026

Spectral Logit Sculpting: Adaptive Low-Rank Logit Transformation for Controlled Text Generation

Jin Li, Zhebo Wang, Tianliang Lu +3

Entropy-based inference methods have gained traction for improving the reliability of Large Language Models (LLMs). However, many existing approaches, such as entropy minimization…

eess.AS2022

Transformer-based Streaming ASR with Cumulative Attention

Mohan Li, Shucong Zhang, Catalin Zorila +1

In this paper, we propose an online attention mechanism, known as cumulative attention (CA), for streaming Transformer-based automatic speech recognition (ASR). Inspired by monoton…

cs.LG2026

Federated Learning with Profile Mapping under Distribution Shifts and Drifts

Mohan Li, Dario Fenoglio, Martin Gjoreski +1

Federated Learning (FL) enables decentralized model training across clients without sharing raw data, but its performance degrades under real-world data heterogeneity. Existing met…

cs.LG2025

FLUX: Efficient Descriptor-Driven Clustered Federated Learning under Arbitrary Distribution Shifts

Dario Fenoglio, Mohan Li, Pietro Barbiero +3

Federated Learning (FL) enables collaborative model training across multiple clients while preserving data privacy. Traditional FL methods often use a global model to fit all clien…

cs.CV2026

Extendable Generalization Self-Supervised Diffusion for Low-Dose CT Reconstruction

Guoquan Wei, Liu Shi, Zekun Zhou +4

Current methods based on deep learning for self-supervised low-dose CT (LDCT) reconstruction, while reducing the dependence on paired data, face the problem of significantly decrea…

cs.CR2026

Neural Honeytrace: Plug&Play Watermarking Framework against Model Extraction Attacks

Yixiao Xu, Binxing Fang, Rui Wang +4

Triggerable watermarking enables model owners to assert ownership against model extraction attacks. However, most existing approaches require additional training, which limits post…

eess.AS2020

Transformer-based Online Speech Recognition with Decoder-end Adaptive Computation Steps

Mohan Li, Catalin Zorila, Rama Doddipatla

Transformer-based end-to-end (E2E) automatic speech recognition (ASR) systems have recently gained wide popularity, and are shown to outperform E2E models based on recurrent struct…

eess.AS2024

WHISMA: A Speech-LLM to Perform Zero-shot Spoken Language Understanding

Mohan Li, Cong-Thanh Do, Simon Keizer +3

Speech large language models (speech-LLMs) integrate speech and text-based foundation models to provide a unified framework for handling a wide range of downstream tasks. In this p…

physics.ins-det2017

Assay of low-background stainless steel by smelting for the neutrino experiment at Jinping

Ghulam Hussain, Zhi Zeng, Chunfa Yao +5

To ensure compliance with the experimental requirement for ultra-low background, in this study the radioactivity of stainless steels manufactured by smelting is thoroughly investig…

physics.ins-det2015

Separation of Scintillation and Cherenkov Lights in Linear Alkyl Benzene

Mohan Li, Ziyi Guo, Minfang Yeh +2

To separate scintillation and Cherenkov lights in water-based liquid scintillator detectors is a desired feature for future neutrino and proton decay researches. Linear alkyl benze…

eess.AS2024

Prompting Whisper for QA-driven Zero-shot End-to-end Spoken Language Understanding

Mohan Li, Simon Keizer, Rama Doddipatla

Zero-shot spoken language understanding (SLU) enables systems to comprehend user utterances in new domains without prior exposure to training data. Recent studies often rely on lar…

cs.CR2025

PREE: Towards Harmless and Adaptive Fingerprint Editing in Large Language Models via Knowledge Prefix Enhancement

Xubin Yue, Zhenhua Xu, Wenpeng Xing +3

Addressing the intellectual property protection challenges in commercial deployment of large language models (LLMs), existing black-box fingerprinting techniques face dual challeng…

eess.AS2023

Non-autoregressive End-to-end Approaches for Joint Automatic Speech Recognition and Spoken Language Understanding

Mohan Li, Rama Doddipatla

This paper presents the use of non-autoregressive (NAR) approaches for joint automatic speech recognition (ASR) and spoken language understanding (SLU) tasks. The proposed NAR syst…

cs.CL2025

SproutBench: A Benchmark for Safe and Ethical Large Language Models for Youth

Wenpeng Xing, Lanyi Wei, Haixiao Hu +5

The rapid proliferation of large language models (LLMs) in applications targeting children and adolescents necessitates a fundamental reassessment of prevailing AI safety framework…

cs.CR2025

Fingerprint Vector: Enabling Scalable and Efficient Model Fingerprint Transfer via Vector Addition

Zhenhua Xu, Qichen Liu, Zhebo Wang +4

Backdoor-based fingerprinting has emerged as an effective technique for tracing the ownership of large language models. However, in real-world deployment scenarios, developers ofte…

physics.ins-det2016

Letter of Intent: Jinping Neutrino Experiment

John F. Beacom, Shaomin Chen, Jianping Cheng +34

Jinping Neutrino Experiment (Jinping) is proposed to significantly improve measurements on solar neutrinos and geoneutrinos in China Jinping Laboratory - a lab with a number of unp…

eess.AS2026

Text-Prompted CLAP: Learning Query-Conditioned Audio Representations via Contrastive Learning

Mohan Li, Rama Doddipatla, Philip C. Woodland

Contrastive Language-Audio Pretraining (CLAP) learns aligned text and audio representations in a shared embedding space. However, independent encoding of each modality limits its a…

cs.CL2026

Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs

Wenpeng Xing, Mohan Li, Bohan Yang +5

Safety-aligned large language models can still be manipulated through white-box interventions that modify their internal representations. We introduce Latent Fusion Jailbreak (LFJ)…

eess.AS2023

Self-regularised Minimum Latency Training for Streaming Transformer-based Speech Recognition

Mohan Li, Rama Doddipatla, Catalin Zorila

This paper proposes a self-regularised minimum latency training (SR-MLT) method for streaming Transformer-based automatic speech recognition (ASR) systems. In previous works, laten…

cs.CL2025

Conditional Multi-Stage Failure Recovery for Embodied Agents

Youmna Farag, Svetlana Stoyanchev, Mohan Li +2

Embodied agents performing complex tasks are susceptible to execution failures, motivating the need for effective failure recovery mechanisms. In this work, we introduce a conditio…

eess.AS2018

End-to-end Speech Recognition with Adaptive Computation Steps

Mohan Li, Min Liu, Masanori Hattori

In this paper, we present Adaptive Computation Steps (ACS) algo-rithm, which enables end-to-end speech recognition models to dy-namically decide how many frames should be processed…

cs.CV2026

SCOUT: Fast Spectral CT Imaging in Ultra LOw-data Regimes via PseUdo-label GeneraTion

Guoquan Wei, Liu Shi, Shaoyu Wang +3

Noise and artifacts during computed tomography (CT) scans are a fundamental challenge affecting disease diagnosis. However, current methods either involve excessively long reconstr…

cs.CL2025

Mitigating Hallucinations in Large Vision-Language Models by Adaptively Constraining Information Flow

Jiaqi Bai, Hongcheng Guo, Zhongyuan Peng +4

Large vision-language models show tremendous potential in understanding visual information through human languages. However, they are prone to suffer from object hallucination, i.e…

cs.LG2026

OrigamiBench: An Interactive Environment to Synthesize Flat-Foldable Origamis

Naaisha Agarwal, Yihan Wu, Yichang Jian +7

Building AI systems that can plan, act, and create in the physical world requires more than pattern recognition. Such systems must understand the causal mechanisms and constraints…

cs.LG2025

Multi-Frequency Federated Learning for Human Activity Recognition Using Head-Worn Sensors

Dario Fenoglio, Mohan Li, Davide Casnici +5

Human Activity Recognition (HAR) benefits various application domains, including health and elderly care. Traditional HAR involves constructing pipelines reliant on centralized use…

cs.CL2026

ICPO: Illocution-Calibrated Policy Optimization for Multi-Turn Conversation

Zhebo Wang, Xiaohu Mu, Zijie Zhou +4

Large Language Models (LLMs) in multi-turn conversations often suffer from a ``lost-in-conversation'' phenomenon, where they struggle to recover from early incorrect assumptions, p…

cs.LG2025

A Survey on Federated Learning in Human Sensing

Mohan Li, Martin Gjoreski, Pietro Barbiero +4

Human Sensing, a field that leverages technology to monitor human activities, psycho-physiological states, and interactions with the environment, enhances our understanding of huma…

cs.CL2022

Multiple-hypothesis RNN-T Loss for Unsupervised Fine-tuning and Self-training of Neural Transducer

Cong-Thanh Do, Mohan Li, Rama Doddipatla

This paper proposes a new approach to perform unsupervised fine-tuning and self-training using unlabeled speech data for recurrent neural network (RNN)-Transducer (RNN-T) end-to-en…

cs.CR2025

Towards Robust and Secure Embodied AI: A Survey on Vulnerabilities and Attacks

Wenpeng Xing, Minghao Li, Mohan Li +1

Embodied AI systems, including robots and autonomous vehicles, are increasingly integrated into real-world applications, where they encounter a range of vulnerabilities stemming fr…

cs.NI2019

Automated Attack and Defense Framework for 5G Security on Physical and Logical Layers

Zhihong Tian, Yanbin Sun, Shen Su +3

The 5th generation (5G) network adopts a great number of revolutionary technologies to fulfill continuously increasing requirements of a variety of applications, including ultra-hi…