activity
20242026
most cited52B to 1T: Lessons Learned via Tele-FLM Series

1 citations · 1 across the 13 of their papers we have counts for

collaborators

13 papers

cs.CL2026

Learning When Not to Listen: Selective Anti-Interference Pretraining for Language Models

Jinchang Zhu, Haowei He, Yi Ding +5

Language models can over-condition on irrelevant preceding text: predictions already supported by local context may still change when distant, unrelated prefix tokens are perturbed…

cs.CL2026

FCPRAG: Fusion-Controller Parametric Retrieval-Augmented Generation for Stable Multi-Passage LoRA Injection

Jinchang Zhu, Jindong Li, Yi Ding +5

Parametric retrieval-augmented generation (PRAG) injects retrieved evidence into a large language model (LLM) through passage-specific LoRA adapters, reducing reliance on long in-c…

cs.CL2026

Dual Attention Residuals

Xingda Yu, Yining Li, Xinzhang Liu +5

Recent work extends Transformer residual pathways along two complementary axes: historical retrieval selects information from earlier depths, whereas multi-stream methods maintain…

cs.CL2026

Hallucinations as Orthogonal Noise: Inference-Time Manifold Alignment via Dynamic Contextual Orthogonalization

Mingkuan Zhao, Wentao Hu, Tianchen Huang +6

Hallucination in Large Language Models (LLMs), characterized by the generation of content inconsistent with contextual facts or logical constraints -- remains a persistent challeng…

cs.LG2026

Awakening Dormant Experts:Counterfactual Routing to Mitigate MoE Hallucinations

Wentao Hu, Yanbo Zhai, Xiaohui Hu +6

Sparse Mixture-of-Experts (MoE) models have achieved remarkable scalability, yet they remain vulnerable to hallucinations, particularly when processing long-tail knowledge. We iden…

cs.LG2026

D-QRELO: Training- and Data-Free Delta Compression for Large Language Models via Quantization and Residual Low-Rank Approximation

Junlin Li, Shuangyong Song, Guodong Du +6

Supervised Fine-Tuning (SFT) accelerates taskspecific large language models (LLMs) development, but the resulting proliferation of finetuned models incurs substantial memory overhe…