collaborators

5 papers

cs.SE2026

From Fragments to Paths: Task-Level Context Recovery for Large Industrial Codebases

Jiawei He, Weisong Sun, Mengyu Shi +4

Large language models have shown strong performance on software engineering (SE) tasks, yet understanding large industrial repositories remains challenging. Existing methods often…

cs.CL2026

SSDAU: Structured Semantic Data Augmentation for Joint Entity and Relation Extraction

Jiawei He, Mengyu Shi, Jiawei Liu +6

Joint Entity and Relation Extraction (JERE) is highly sensitive to training data quality, making data augmentation a natural way to improve generalization. However, existing augmen…

cs.SE2026

ProcCtrlBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents

Jiawei He, Jie Jia, Chenbo Liu +4

Existing benchmarks for LLM coding agents primarily evaluate final outcomes. While useful for measuring overall capability, these metrics provide limited visibility and often miss…

q-bio.OT2025

Retrieval-Augmented Generation in Biomedicine: A Survey of Technologies, Datasets, and Clinical Applications

Jiawei He, Boya Zhang, Hossein Rouhizadeh +6

Large language models (LLMs) in biomedicine face a fundamental conflict between static parameter knowledge and the dynamic nature of clinical evidence. Retrieval-Augmented Generati…

cs.CL2025

ICA-RAG: Information Completeness Guided Adaptive Retrieval-Augmented Generation for Disease Diagnosis

Jiawei He, Mingyi Jia, Zhihao Jia +3

Retrieval-Augmented Large Language Models (LLMs), which integrate external knowledge, have shown remarkable performance in medical domains, including clinical diagnosis. However, e…