4 papers
CEC-Zero: Zero-Supervision Character Error Correction with Self-Generated Rewards
Zhiming Lin, Kai Zhao, Sophie Zhang +2
Large-scale Chinese spelling correction (CSC) remains critical for real-world text processing, yet existing LLMs and supervised methods lack robustness to novel errors and rely on…
From Points to Coalitions: Hierarchical Contrastive Shapley Values for Prioritizing Data Samples
Canran Xiao, Jiabao Dou, Zhiming Lin +2
How should we quantify the value of each training example when datasets are large, heterogeneous, and geometrically structured? Classical Data-Shapley answers in principle, but its…
Beyond More Context: Retrieval Diversity Boosts Multi-Turn Intent Understanding
Zhiming Lin
Multi turn intent understanding is central to task oriented chatbots, yet real deployments face tight token budgets and noisy contexts, and most retrieval pipelines emphasize relev…
CEC-Zero: Chinese Error Correction Solution Based on LLM
Sophie Zhang, Zhiming Lin
Recent advancements in large language models (LLMs) demonstrate exceptional Chinese text processing capabilities, particularly in Chinese Spelling Correction (CSC). While LLMs outp…