5 papers
From Fragments to Paths: Task-Level Context Recovery for Large Industrial Codebases
Jiawei He, Weisong Sun, Mengyu Shi +4
Large language models have shown strong performance on software engineering (SE) tasks, yet understanding large industrial repositories remains challenging. Existing methods often…
SSDAU: Structured Semantic Data Augmentation for Joint Entity and Relation Extraction
Jiawei He, Mengyu Shi, Jiawei Liu +6
Joint Entity and Relation Extraction (JERE) is highly sensitive to training data quality, making data augmentation a natural way to improve generalization. However, existing augmen…
ProcCtrlBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents
Jiawei He, Jie Jia, Chenbo Liu +4
Existing benchmarks for LLM coding agents primarily evaluate final outcomes. While useful for measuring overall capability, these metrics provide limited visibility and often miss…
Retrieval-Augmented Generation in Biomedicine: A Survey of Technologies, Datasets, and Clinical Applications
Jiawei He, Boya Zhang, Hossein Rouhizadeh +6
Large language models (LLMs) in biomedicine face a fundamental conflict between static parameter knowledge and the dynamic nature of clinical evidence. Retrieval-Augmented Generati…
ICA-RAG: Information Completeness Guided Adaptive Retrieval-Augmented Generation for Disease Diagnosis
Jiawei He, Mingyi Jia, Zhihao Jia +3
Retrieval-Augmented Large Language Models (LLMs), which integrate external knowledge, have shown remarkable performance in medical domains, including clinical diagnosis. However, e…