16 papers
Calibrating Post-Training Feature Shifts for LLM Data Contamination Detection
Zhen Yang, Mengqi Wang, Gengda Zhao +3
Large language models (LLMs) are trained on massive and largely undisclosed corpora that may contain copyrighted or privacy-sensitive content. Data contamination detection (DCD) th…
Interpretable Unsupervised Community Detection with LLM-Symbolized Structured Processes
Aoting Zeng, Kai Wang, Jianwei Wang +3
Community detection is a fundamental task in graph analytics that aims to identify cohesive groups of entities with similar behaviors or interests. Classic objective-driven methods…
Interpretable Column Annotation with LLM-Symbolized Decision Process Materialization
Mengqi Wang, Jianwei Wang, Qing Liu +5
Column annotation (CA), including column type annotation (CTA) and column property annotation (CPA), aims to identify the meanings of table columns and the semantic relationships a…
Collaborative Large and Small Language Models for Accurate and Scalable Data Repair
Qian Chen, Jianwei Wang, Wenjie Zhang
We study the problem of data repair, a key task in data cleaning that corrects erroneous entries in raw datasets to improve overall data quality. Although recent data-driven method…
EvoSQL: Memory-Augmented Critic-Generator Co-Evolution for Text-to-SQL
Jiawei Zhou, Jianwei Wang, Chenyu Zhou +3
Text-to-SQL has advanced rapidly with large language models, but complex database queries still require reasoning beyond one-shot generation, including multi-step decomposition, ex…
Multi-Perspective Evidence Synthesis and Reasoning for Unsupervised Multimodal Entity Linking
Mo Zhou, Jianwei Wang, Kai Wang +3
Multimodal Entity Linking (MEL) is a fundamental task in data management that maps ambiguous mentions with diverse modalities to the multimodal entities in a knowledge base. Howeve…