30 citations · 49 across the 4 of their papers we have counts for
6 papers
DOM-LM: Learning Generalizable Representations for HTML Documents
Xiang Deng, Prashant Shiralkar, Colin Lockard +2
HTML documents are an important medium for disseminating information on the Web for human consumption. An HTML document presents information in multiple text formats including unst…
TCN: Table Convolutional Network for Web Table Interpretation
Daheng Wang, Prashant Shiralkar, Colin Lockard +3
Information extraction from semi-structured webpages provides valuable long-tailed facts for augmenting knowledge graph. Relational Web tables are a critical component containing a…
ZeroShotCeres: Zero-Shot Relation Extraction from Semi-Structured Webpages
Colin Lockard, Prashant Shiralkar, Xin Luna Dong +1
In many documents, such as semi-structured webpages, textual semantics are augmented with additional information conveyed using visual elements including layout, font size, and col…
OpenKI: Integrating Open Information Extraction and Knowledge Bases with Relation Inference
Dongxu Zhang, Subhabrata Mukherjee, Colin Lockard +2
In this paper, we consider advancing web-scale knowledge extraction and alignment by integrating OpenIE extractions in the form of (subject, predicate, object) triples with Knowled…
Semi-Supervised Event Extraction with Paraphrase Clusters
James Ferguson, Colin Lockard, Daniel S. Weld +1
Supervised event extraction systems are limited in their accuracy due to the lack of available training data. We present a method for self-training event extraction systems by boot…
CERES: Distantly Supervised Relation Extraction from the Semi-Structured Web
Colin Lockard, Xin Luna Dong, Arash Einolghozati +1
The web contains countless semi-structured websites, which can be a rich source of information for populating knowledge bases. Existing methods for extracting relations from the DO…