7 papers · 1 filter
Calibrating Post-Training Feature Shifts for LLM Data Contamination Detection
Zhen Yang, Mengqi Wang, Gengda Zhao +3
Large language models (LLMs) are trained on massive and largely undisclosed corpora that may contain copyrighted or privacy-sensitive content. Data contamination detection (DCD) th…
Interpretable Column Annotation with LLM-Symbolized Decision Process Materialization
Mengqi Wang, Jianwei Wang, Qing Liu +5
Column annotation (CA), including column type annotation (CTA) and column property annotation (CPA), aims to identify the meanings of table columns and the semantic relationships a…
Multi-Perspective Evidence Synthesis and Reasoning for Unsupervised Multimodal Entity Linking
Mo Zhou, Jianwei Wang, Kai Wang +3
Multimodal Entity Linking (MEL) is a fundamental task in data management that maps ambiguous mentions with diverse modalities to the multimodal entities in a knowledge base. Howeve…
HyperJoin: LLM-augmented Hypergraph Link Prediction for Joinable Table Discovery
Shiyuan Liu, Jianwei Wang, Xuemin Lin +3
As a pivotal task in data lake management, joinable table discovery has attracted widespread interest. While existing language model-based methods achieve remarkable performance by…
Ensembling LLM-Induced Decision Trees for Explainable and Robust Error Detection
Mengqi Wang, Jianwei Wang, Qing Liu +5
Error detection (ED), which aims to identify incorrect or inconsistent cell values in tabular data, is important for ensuring data quality. Recent state-of-the-art ED methods lever…
Large Language Models Meet Text-Attributed Graphs: A Survey of Integration Frameworks and Applications
Guangxin Su, Hanchen Wang, Jianwei Wang +3
Large Language Models (LLMs) have achieved remarkable success in natural language processing through strong semantic understanding and generation. However, their black-box nature l…