activity
20242026
most citedNoise-Resilient Point-wise Anomaly Detection in Time Series Using Weak Segment Labels

3 citations · 3 across the 4 of their papers we have counts for

collaborators

6 papers

cs.AI2026

Small-Margin Preferences Still Matter-If You Train Them Right

Jinlong Pang, Zhaowei Zhu, Na Di +4

Preference optimization methods such as DPO align large language models (LLMs) using paired comparisons, but their effectiveness can be highly sensitive to the quality and difficul…

cs.AI2025

Evaluating LLM-Contaminated Crowdsourcing Data Without Ground Truth

Yichi Zhang, Jinlong Pang, Zhaowei Zhu +1

The recent success of generative AI highlights the crucial role of high-quality human feedback in building trustworthy AI systems. However, the increasing use of large language mod…

cs.CL2025

Token Cleaning: Fine-Grained Data Selection for LLM Supervised Fine-Tuning

Jinlong Pang, Na Di, Zhaowei Zhu +4

Recent studies show that in supervised fine-tuning (SFT) of large language models (LLMs), data quality matters more than quantity. While most data cleaning methods concentrate on f…

cs.LG20253 cited

Noise-Resilient Point-wise Anomaly Detection in Time Series Using Weak Segment Labels

Yaxuan Wang, Hao Cheng, Jing Xiong +6

Detecting anomalies in temporal data has gained significant attention across various real-world applications, aiming to identify unusual events and mitigate potential hazards. In p…

cs.LG2024

Reassessing Layer Pruning in LLMs: New Insights and Methods

Yao Lu, Hao Cheng, Yujie Fang +6

Although large language models (LLMs) have achieved remarkable success across various domains, their considerable scale necessitates substantial computational resources, posing sig…

cs.CL2024

Improving Data Efficiency via Curating LLM-Driven Rating Systems

Jinlong Pang, Jiaheng Wei, Ankit Parag Shah +6

Instruction tuning is critical for adapting large language models (LLMs) to downstream tasks, and recent studies have demonstrated that small amounts of human-curated data can outp…