Publications (34)
DiaLoc: An Iterative Approach to Embodied Dialog Localization
Chao Zhang, Mohan Li, Ignas Budvytis +1
Multimodal learning has advanced the performance for many vision-language tasks. However, most existing works in embodied dialog research focus on navigation and leave the localiza…
Head-synchronous Decoding for Transformer-based Streaming ASR
Mohan Li, Catalin Zorila, Rama Doddipatla
Online Transformer-based automatic speech recognition (ASR) systems have been extensively studied due to the increasing demand for streaming applications. Recently proposed Decoder…
Economic Forces in Stock Returns
Yue Chen, Mohan Li
When analyzing the components influencing the stock prices, it is commonly believed that economic activities play an important role. More specifically, asset prices are more sensit…
FIDES: Faithful Inference via Deep Evidence Signals for Retrieval-Memory Conflict in RAG
Zhe Yu, Wenpeng Xing, Tiancheng Zhao +3
When retrieved evidence contradicts parametric memory, language models frequently ignore context and default to memorized priors -- a failure that undermines the core purpose of re…
Spectral Logit Sculpting: Adaptive Low-Rank Logit Transformation for Controlled Text Generation
Jin Li, Zhebo Wang, Tianliang Lu +3
Entropy-based inference methods have gained traction for improving the reliability of Large Language Models (LLMs). However, many existing approaches, such as entropy minimization…
Transformer-based Streaming ASR with Cumulative Attention
Mohan Li, Shucong Zhang, Catalin Zorila +1
In this paper, we propose an online attention mechanism, known as cumulative attention (CA), for streaming Transformer-based automatic speech recognition (ASR). Inspired by monoton…
Federated Learning with Profile Mapping under Distribution Shifts and Drifts
Mohan Li, Dario Fenoglio, Martin Gjoreski +1
Federated Learning (FL) enables decentralized model training across clients without sharing raw data, but its performance degrades under real-world data heterogeneity. Existing met…
FLUX: Efficient Descriptor-Driven Clustered Federated Learning under Arbitrary Distribution Shifts
Dario Fenoglio, Mohan Li, Pietro Barbiero +3
Federated Learning (FL) enables collaborative model training across multiple clients while preserving data privacy. Traditional FL methods often use a global model to fit all clien…
Extendable Generalization Self-Supervised Diffusion for Low-Dose CT Reconstruction
Guoquan Wei, Liu Shi, Zekun Zhou +4
Current methods based on deep learning for self-supervised low-dose CT (LDCT) reconstruction, while reducing the dependence on paired data, face the problem of significantly decrea…
Neural Honeytrace: Plug&Play Watermarking Framework against Model Extraction Attacks
Yixiao Xu, Binxing Fang, Rui Wang +4
Triggerable watermarking enables model owners to assert ownership against model extraction attacks. However, most existing approaches require additional training, which limits post…
Transformer-based Online Speech Recognition with Decoder-end Adaptive Computation Steps
Mohan Li, Catalin Zorila, Rama Doddipatla
Transformer-based end-to-end (E2E) automatic speech recognition (ASR) systems have recently gained wide popularity, and are shown to outperform E2E models based on recurrent struct…
WHISMA: A Speech-LLM to Perform Zero-shot Spoken Language Understanding
Mohan Li, Cong-Thanh Do, Simon Keizer +3
Speech large language models (speech-LLMs) integrate speech and text-based foundation models to provide a unified framework for handling a wide range of downstream tasks. In this p…
Assay of low-background stainless steel by smelting for the neutrino experiment at Jinping
Ghulam Hussain, Zhi Zeng, Chunfa Yao +5
To ensure compliance with the experimental requirement for ultra-low background, in this study the radioactivity of stainless steels manufactured by smelting is thoroughly investig…
Separation of Scintillation and Cherenkov Lights in Linear Alkyl Benzene
Mohan Li, Ziyi Guo, Minfang Yeh +2
To separate scintillation and Cherenkov lights in water-based liquid scintillator detectors is a desired feature for future neutrino and proton decay researches. Linear alkyl benze…
Prompting Whisper for QA-driven Zero-shot End-to-end Spoken Language Understanding
Mohan Li, Simon Keizer, Rama Doddipatla
Zero-shot spoken language understanding (SLU) enables systems to comprehend user utterances in new domains without prior exposure to training data. Recent studies often rely on lar…
PREE: Towards Harmless and Adaptive Fingerprint Editing in Large Language Models via Knowledge Prefix Enhancement
Xubin Yue, Zhenhua Xu, Wenpeng Xing +3
Addressing the intellectual property protection challenges in commercial deployment of large language models (LLMs), existing black-box fingerprinting techniques face dual challeng…
Non-autoregressive End-to-end Approaches for Joint Automatic Speech Recognition and Spoken Language Understanding
Mohan Li, Rama Doddipatla
This paper presents the use of non-autoregressive (NAR) approaches for joint automatic speech recognition (ASR) and spoken language understanding (SLU) tasks. The proposed NAR syst…
SproutBench: A Benchmark for Safe and Ethical Large Language Models for Youth
Wenpeng Xing, Lanyi Wei, Haixiao Hu +5
The rapid proliferation of large language models (LLMs) in applications targeting children and adolescents necessitates a fundamental reassessment of prevailing AI safety framework…
Fingerprint Vector: Enabling Scalable and Efficient Model Fingerprint Transfer via Vector Addition
Zhenhua Xu, Qichen Liu, Zhebo Wang +4
Backdoor-based fingerprinting has emerged as an effective technique for tracing the ownership of large language models. However, in real-world deployment scenarios, developers ofte…
Letter of Intent: Jinping Neutrino Experiment
John F. Beacom, Shaomin Chen, Jianping Cheng +34
Jinping Neutrino Experiment (Jinping) is proposed to significantly improve measurements on solar neutrinos and geoneutrinos in China Jinping Laboratory - a lab with a number of unp…
Text-Prompted CLAP: Learning Query-Conditioned Audio Representations via Contrastive Learning
Mohan Li, Rama Doddipatla, Philip C. Woodland
Contrastive Language-Audio Pretraining (CLAP) learns aligned text and audio representations in a shared embedding space. However, independent encoding of each modality limits its a…
Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs
Wenpeng Xing, Mohan Li, Bohan Yang +5
Safety-aligned large language models can still be manipulated through white-box interventions that modify their internal representations. We introduce Latent Fusion Jailbreak (LFJ)…
Self-regularised Minimum Latency Training for Streaming Transformer-based Speech Recognition
Mohan Li, Rama Doddipatla, Catalin Zorila
This paper proposes a self-regularised minimum latency training (SR-MLT) method for streaming Transformer-based automatic speech recognition (ASR) systems. In previous works, laten…
Conditional Multi-Stage Failure Recovery for Embodied Agents
Youmna Farag, Svetlana Stoyanchev, Mohan Li +2
Embodied agents performing complex tasks are susceptible to execution failures, motivating the need for effective failure recovery mechanisms. In this work, we introduce a conditio…
End-to-end Speech Recognition with Adaptive Computation Steps
Mohan Li, Min Liu, Masanori Hattori
In this paper, we present Adaptive Computation Steps (ACS) algo-rithm, which enables end-to-end speech recognition models to dy-namically decide how many frames should be processed…
SCOUT: Fast Spectral CT Imaging in Ultra LOw-data Regimes via PseUdo-label GeneraTion
Guoquan Wei, Liu Shi, Shaoyu Wang +3
Noise and artifacts during computed tomography (CT) scans are a fundamental challenge affecting disease diagnosis. However, current methods either involve excessively long reconstr…
Mitigating Hallucinations in Large Vision-Language Models by Adaptively Constraining Information Flow
Jiaqi Bai, Hongcheng Guo, Zhongyuan Peng +4
Large vision-language models show tremendous potential in understanding visual information through human languages. However, they are prone to suffer from object hallucination, i.e…
OrigamiBench: An Interactive Environment to Synthesize Flat-Foldable Origamis
Naaisha Agarwal, Yihan Wu, Yichang Jian +7
Building AI systems that can plan, act, and create in the physical world requires more than pattern recognition. Such systems must understand the causal mechanisms and constraints…
Multi-Frequency Federated Learning for Human Activity Recognition Using Head-Worn Sensors
Dario Fenoglio, Mohan Li, Davide Casnici +5
Human Activity Recognition (HAR) benefits various application domains, including health and elderly care. Traditional HAR involves constructing pipelines reliant on centralized use…
ICPO: Illocution-Calibrated Policy Optimization for Multi-Turn Conversation
Zhebo Wang, Xiaohu Mu, Zijie Zhou +4
Large Language Models (LLMs) in multi-turn conversations often suffer from a ``lost-in-conversation'' phenomenon, where they struggle to recover from early incorrect assumptions, p…
A Survey on Federated Learning in Human Sensing
Mohan Li, Martin Gjoreski, Pietro Barbiero +4
Human Sensing, a field that leverages technology to monitor human activities, psycho-physiological states, and interactions with the environment, enhances our understanding of huma…
Multiple-hypothesis RNN-T Loss for Unsupervised Fine-tuning and Self-training of Neural Transducer
Cong-Thanh Do, Mohan Li, Rama Doddipatla
This paper proposes a new approach to perform unsupervised fine-tuning and self-training using unlabeled speech data for recurrent neural network (RNN)-Transducer (RNN-T) end-to-en…
Towards Robust and Secure Embodied AI: A Survey on Vulnerabilities and Attacks
Wenpeng Xing, Minghao Li, Mohan Li +1
Embodied AI systems, including robots and autonomous vehicles, are increasingly integrated into real-world applications, where they encounter a range of vulnerabilities stemming fr…
Automated Attack and Defense Framework for 5G Security on Physical and Logical Layers
Zhihong Tian, Yanbin Sun, Shen Su +3
The 5th generation (5G) network adopts a great number of revolutionary technologies to fulfill continuously increasing requirements of a variety of applications, including ultra-hi…