Publications (31)
ERNIE 5.0 Technical Report
Haifeng Wang, Hua Wu, Tian Wu +432
In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…
HERO: History-Enriched Rollout Training for Long-Horizon Autoregressive Neural Operators
Jiaquan Zhang, Shuxu Chen, Haifan Meng +6
Neural operators provide fast surrogates for time-dependent partial differential equations (PDEs) by applying a learned evolution operator recursively to its own predictions, but t…
Mistletoe: Stealthy Acceleration-Collapse Attacks on Speculative Decoding
Shuoyang Sun, Chang Dai, Hao Fang +6
Speculative decoding has become a widely adopted technique for accelerating large language model (LLM) inference by drafting multiple candidate tokens and verifying them with a tar…
Personalized Fashion Recommendation with Image Attributes and Aesthetics Assessment
Chongxian Chen, Fan Mo, Xin Fan +1
Personalized fashion recommendation is a difficult task because 1) the decisions are highly correlated with users' aesthetic appetite, which previous work frequently overlooks, and…
Machine Learning with Confidential Computing: A Systematization of Knowledge
Fan Mo, Zahra Tarkhani, Hamed Haddadi
Privacy and security challenges in Machine Learning (ML) have become increasingly severe, along with ML's pervasive development and the recent demonstration of large attack surface…
DeepRT: deep learning for peptide retention time prediction in proteomics
Chunwei Ma, Zhiyong Zhu, Jun Ye +8
Accurate predictions of peptide retention times (RT) in liquid chromatography have many applications in mass spectrometry-based proteomics. Herein, we present DeepRT, a deep learni…
Towards Characterizing and Limiting Information Exposure in DNN Layers
Fan Mo, Ali Shahin Shamsabadi, Kleomenis Katevas +2
Pre-trained Deep Neural Network (DNN) models are increasingly used in smartphones and other user devices to enable prediction services, leading to potential disclosures of (sensiti…
Private delegated computations using strong isolation
Mathias Brossard, Guilhem Bryant, Basma El Gaabouri +12
Sensitive computations are now routinely delegated to third-parties. In response, Confidential Computing technologies are being introduced to microprocessors, offering a protected…
Weight-importance sparse training in keyword spotting
Sihao Xue, Zhenyi Ying, Fan Mo +2
Large size models are implemented in recently ASR system to deal with complex speech recognition problems. The num- ber of parameters in these models makes them hard to deploy, esp…
Connectivity-Guided Sparsification of 2-FWL GNNs: Preserving Full Expressivity with Improved Efficiency
Rongqin Chen, Fan Mo, Pak Lon Ip +4
Higher-order Graph Neural Networks (HOGNNs) based on the 2-FWL test achieve superior expressivity by modeling 2- and 3-node interactions, but at computational co…
MemVenom: Triggered Poisoning of Multimodal Memories in Web Agents
Yv Zhang, Hao Sun, Hao Fang +5
External memory has become a core component of modern web agents, enabling long-horizon reasoning through the retrieval of past experiences. However, this paradigm introduces a cri…
KG4Diagnosis: A Hierarchical Multi-Agent LLM Framework with Knowledge Graph Enhancement for Medical Diagnosis
Kaiwen Zuo, Yirui Jiang, Fan Mo +1
Integrating Large Language Models (LLMs) in healthcare diagnosis demands systematic frameworks that can handle complex medical scenarios while maintaining specialized expertise. We…
Enhancing Efficiency in Multidevice Federated Learning through Data Selection
Fan Mo, Mohammad Malekzadeh, Soumyajit Chatterjee +2
Ubiquitous wearable and mobile devices provide access to a diverse set of data. However, the mobility demand for our devices naturally imposes constraints on their computational an…
DarkneTZ: Towards Model Privacy at the Edge using Trusted Execution Environments
Fan Mo, Ali Shahin Shamsabadi, Kleomenis Katevas +4
We present DarkneTZ, a framework that uses an edge device's Trusted Execution Environment (TEE) in conjunction with model partitioning to limit the attack surface against Deep Neur…
Round Attention: A Novel Round-Level Attention Mechanism to Accelerate LLM Inference
Yaohua Tang, Zhicheng Hu, Kun Cheng +4
The increasing context window size in large language models (LLMs) has improved their ability to handle complex, long-text tasks. However, as the conversation rounds continue, it i…
Towards Battery-Free Machine Learning and Inference in Underwater Environments
Yuchen Zhao, Sayed Saad Afzal, Waleed Akbar +5
This paper is motivated by a simple question: Can we design and build battery-free devices capable of machine learning and inference in underwater environments? An affirmative answ…
Quantifying and Localizing Usable Information Leakage from Neural Network Gradients
Fan Mo, Anastasia Borovykh, Mohammad Malekzadeh +3
In collaborative learning, clients keep their data private and communicate only the computed gradients of the deep neural network being trained on their local data. Several recent…
PPFL: Privacy-preserving Federated Learning with Trusted Execution Environments
Fan Mo, Hamed Haddadi, Kleomenis Katevas +3
We propose and implement a Privacy-preserving Federated Learning () framework for mobile systems to limit privacy leakages in federated learning. Leveraging the widespread pr…
PORTool: Importance-Aware Policy Optimization with Rewarded Tree for Multi-Tool-Integrated Reasoning
Feijie Wu, Weiwu Zhu, Yuxiang Zhang +5
Multi-tool-integrated reasoning enables LLM-empowered tool-use agents to solve complex tasks by interleaving natural-language reasoning with calls to external tools. However, train…
TF-SNO: Time-Frequency Gated Spectral Neural Operators for Learning Non-Stationary Partial Differential Equations
Yitian Zhou, Chaoning Zhang, Zhenzhen Huang +8
Non-stationary partial differential equations (PDEs) arise throughout scientific computing, where the dominant frequency content and energy distribution can drift over time. While…
Layer-wise Characterization of Latent Information Leakage in Federated Learning
Fan Mo, Anastasia Borovykh, Mohammad Malekzadeh +2
Training deep neural networks via federated learning allows clients to share, instead of the original data, only the model trained on their data. Prior work has demonstrated that i…
Reference-Free Image Quality Assessment for Virtual Try-On via Human Feedback
Yuki Hirakawa, Takashi Wada, Ryotaro Shimizu +6
As virtual try-on (VTON) systems become increasingly important in fashion e-commerce, there is a growing need for reliable reference-free evaluation methods, since ground-truth ima…
ELAN4D: Embodiment-Centric 4D Supervision for Vision-Language-Action Models via Plug-and-Play Adaptation
Zeyuan He, Bowen Yang, Zhirui Fang +9
Vision-Language-Action (VLA) models have shown promise for robotic manipulation, yet most existing policies operate reactively by directly regressing actions from current observati…
Geometric Neural Operators via Lie Group-Constrained Latent Dynamics
Jiaquan Zhang, Fachrina Dewi Puspitasari, Songbo Zhang +7
Neural operators offer an effective framework for learning solutions of partial differential equations for many physical systems in a resolution-invariant and data-driven manner. E…
Retention Time of Peptides in Liquid Chromatography Is Well Estimated upon Deep Transfer Learning
Chunwei Ma, Zhiyong Zhu, Jun Ye +7
A fully automatic prediction for peptide retention time (RT) in liquid chromatography (LC), termed as DeepRT, was developed using deep learning approach, an ensemble of Residual Ne…
FlexMoE: One-for-All Nested Intra-Expert Pruning for MoE Language Models
Fan Mo, Yuxuan Han, Geng Zhang +2
Mixture-of-Experts (MoE) language models scale model ability with sparsely activated experts, making this architecture a standard recipe for modern large models. However, sparse ac…
How to make Medical AI Systems safer? Simulating Vulnerabilities, and Threats in Multimodal Medical RAG System
Kaiwen Zuo, Zelin Liu, Raman Dutt +4
Large Vision-Language Models (LVLMs) augmented with Retrieval-Augmented Generation (RAG) are increasingly employed in medical AI to enhance factual grounding through external clini…
Multi-Robot Open Adaptive Teaming Across Unseen Environments, Partners, and Scales
Yang Li, Feng Xue, Fan Mo +6
Deploying robot teams in the real world requires simultaneous adaptation to unseen environments, unknown partners, and varying team sizes, yet existing approaches often address the…
Lightweight LLM Agent Memory with Small Language Models
Jiaquan Zhang, Chaoning Zhang, Shuxu Chen +9
Although LLM agents can leverage tools for complex tasks, they still need memory to maintain cross-turn consistency and accumulate reusable information in long-horizon interactions…
Mitigating Hallucinations in Large Language Models via Self-Refinement-Enhanced Knowledge Retrieval
Mengjia Niu, Hao Li, Jie Shi +2
Large language models (LLMs) have demonstrated remarkable capabilities across various domains, although their susceptibility to hallucination poses significant challenges for their…
Data-Efficient Massive Tool Retrieval: A Reinforcement Learning Approach for Query-Tool Alignment with Language Models
Yuxiang Zhang, Xin Fan, Junjie Wang +4
Recent advancements in large language models (LLMs) integrated with external tools and APIs have successfully addressed complex tasks by using in-context learning or fine-tuning. D…