6 papers
AutoEmpirical: LLM-Based Automated Research for Empirical Software Fault Analysis
Jiongchi Yu, Weipeng Jiang, Xiaoyu Zhang +3
Understanding software faults is essential for empirical research in software development and maintenance. However, traditional fault analysis, while valuable, typically involves m…
Mitigating Stylistic Biases of Machine Translation Systems via Monolingual Corpora Only
Xuanqi Gao, Weipeng Jiang, Juan Zhai +4
The advent of neural machine translation (NMT) has revolutionized cross-lingual communication, yet preserving stylistic nuances remains a significant challenge. While existing appr…
The Foundation Cracks: A Comprehensive Study on Bugs and Testing Practices in LLM Libraries
Weipeng Jiang, Xiaoyu Zhang, Xiaofei Xie +4
Large Language Model (LLM) libraries have emerged as the foundational infrastructure powering today's AI revolution, serving as the backbone for LLM deployment, inference optimizat…
Holistic Audit Dataset Generation for LLM Unlearning via Knowledge Graph Traversal and Redundancy Removal
Weipeng Jiang, Juan Zhai, Shiqing Ma +4
In recent years, Large Language Models (LLMs) have faced increasing demands to selectively remove sensitive information, protect privacy, and comply with copyright regulations thro…
The Invisible Hand: Unveiling Provider Bias in Large Language Models for Code Generation
Xiaoyu Zhang, Juan Zhai, Shiqing Ma +5
Large Language Models (LLMs) have emerged as the new recommendation engines, surpassing traditional methods in both capability and scope, particularly in code generation. In this p…
Speculative Coreset Selection for Task-Specific Fine-tuning
Xiaoyu Zhang, Juan Zhai, Shiqing Ma +4
Task-specific fine-tuning is essential for the deployment of large language models (LLMs), but it requires significant computational resources and time. Existing solutions have pro…