collaborators

9 papers

cs.CL2026

DiSRouter: Distributed Self-Routing for LLM Selections

Hang Zheng, Hongshen Xu, Yongkai Lin +3

The proliferation of Large Language Models (LLMs) has created a diverse ecosystem of models with highly varying performance and costs, necessitating effective query routing to bala…

cs.CL2025

Enhancing LLM Reliability via Explicit Knowledge Boundary Modeling

Hang Zheng, Hongshen Xu, Yuncong Liu +3

Large language models (LLMs) are prone to hallucination stemming from misaligned self-awareness, particularly when processing queries exceeding their knowledge boundaries. While ex…

cs.CL2025

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity

Da Ma, Lu Chen, Situo Zhang +8

The rapid expansion of context window sizes in Large Language Models~(LLMs) has enabled them to tackle increasingly complex tasks involving lengthy documents. However, this progres…

cs.CL2025

Developing ChemDFM as a large language foundation model for chemistry

Zihan Zhao, Da Ma, Lu Chen +11

Artificial intelligence (AI) has played an increasingly important role in chemical research. However, most models currently used in chemistry are specialist models that require tra…

cs.CL2025

Reducing Tool Hallucination via Reliability Alignment

Hongshen Xu, Zichen Zhu, Lei Pan +6

Large Language Models (LLMs) have expanded their capabilities beyond language generation to interact with external tools, enabling automation and real-world applications. However,…

cs.LG2025

Task-Specific Data Selection for Instruction Tuning via Monosemantic Neuronal Activations

Da Ma, Gonghu Shang, Zhi Chen +6

Instruction tuning improves the ability of large language models (LLMs) to follow diverse human instructions, but achieving strong performance on specific target tasks remains chal…