most citedMoE: Optimizing Collaborative Inference for Edge Large Language Models

1 citations · 1 across the 4 of their papers we have counts for

collaborators

12 papers

cs.CR2026

Mind the Hook: Source-Level Auditing of Privacy Defenses in Retrieval-Augmented Generation

Yanhang Li, Zhichao Fan, Zexin Zhuang

Black-box privacy scores for retrieval-augmented generation (RAG) are difficult to interpret unless the audited defense's active pipeline hook is known. We propose an active-path a…

cs.SE2026

The Moving Target: A Longitudinal Audit of Trust-Benchmark Score Drift Across Open-Source Chat LLM Release Lines

Zhichao Fan, Yanhang Li, Zexin Zhuang +2

Trust-benchmark scores reported on a chat-LLM release line are often carried across several checkpoints of the same line, as if the underlying model had not shifted between release…

cs.CL2026

Exploratory As-Analyzed No-Detection of Culturally-Marked Predicate-Triggered PII Amplification in a Synthetic-English RAG Probe: A Predicate-Resource-Confounded Audit

Yanhang Li, Zhichao Fan, Zexin Zhuang

We ask whether stereotype-loaded queries about culturally marked people leak more personal information from a retrieval-augmented generation (RAG) system than otherwise-equivalent…

cs.CL2026

Where Privacy Risk Lives in English-Source Multilingual RAG: A Stage-Decomposed Audit Across Five Query Languages

Yanhang Li, Zhichao Fan, Zexin Zhuang

A common assumption holds that switching to a non-English language makes a multilingual RAG system easier to attack for personal information. We test this on an English-source synt…

cs.LG2026

Auditing the Audit: Five Failure Modes in Benchmark-Validity Audits

Yanhang Li, Zhichao Fan, Zexin Zhuang

Governance frameworks ask AI providers and auditors for documented evaluation evidence, and perturbation-based construct-validity audits are a common form of that evidence. We argu…

cs.LG2026

Do Time Series Foundation Model Benchmarks Hide Regime-Dependent Failures? Evidence from Traffic Speed Forecasting

Yingshuo Wang, Xian Sun, Lingdong Kong +4

Standard benchmarks evaluate time series foundation models (TSFMs) using aggregate metrics, but these can mask severe failures in critical operating regimes. We introduce regime-st…