activity
20242026
collaborators

8 papers

cs.CL2026

Robust Reasoning via Dynamic Token Selection for Distribution-Aligned Self-Distillation

Ruiqi Zhang, Lingxiang Wang, Hainan Zhang Zhiming Zheng

Self-distillation improves learning efficiency by rewriting reference answers as training data that better matches the model's own distribution. However, reference answers also int…

cs.CL2026

From Unfamiliar to Familiar: Detecting Pre-training Data via Gradient Deviations in Large Language Models

Ruiqi Zhang, Lingxiang Wang, Hainan Zhang +2

Pre-training data detection for LLMs is essential for addressing copyright concerns and mitigating benchmark contamination. Existing methods mainly focus on the likelihood-based st…

cs.CL2026

LocalSUG: City-Preference-Enhanced LLM for Query Suggestion in Local-Life Services

Jinwen Chen, Shiwen Zhang, Shuai Gong +6

In local-life service platforms, query suggestion reduces user effort by generating candidate queries from input prefixes. Traditional multi-stage systems rely heavily on historica…

cs.LG2025

Parameter Importance-Driven Continual Learning for Foundation Models

Lingxiang Wang, Hainan Zhang, Zhiming Zheng

Domain-specific post-training often causes catastrophic forgetting, making foundation models lose their general reasoning ability and limiting their adaptability to dynamic real-wo…

cs.CL2025

FedDTRE: Federated Dialogue Generation Models Powered by Trustworthiness Evaluation

Shule Lu, Lingxiang Wang, Sijia Wen +2

With the rapid development of artificial intelligence, dialogue systems have become a prominent form of human-computer interaction. However, traditional centralized or fully local…

cs.IR2025

Beyond the Surface: A Solution-Aware Retrieval Model for Competition-level Code Generation

Shiwen Zhang, Lingxiang Wang, Hainan Zhang +3

In competitive programming task, problem statements are often embedded within elaborate narrative backgrounds, requiring deep understanding of the underlying solutions to successfu…