most citedScanBank: A Benchmark Dataset for Figure Extraction from Scanned Electronic Theses and Dissertations

1 citations · 1 across the 4 of their papers we have counts for

collaborators

5 papers

cs.DL2021

Automatic Metadata Extraction Incorporating Visual Features from Scanned Electronic Theses and Dissertations

Muntabir Hasan Choudhury, Himarsha R. Jayanetti, Jian Wu +2

Electronic Theses and Dissertations (ETDs) contain domain knowledge that can be used for many digital library tasks, such as analyzing citation networks and predicting research tre…

cs.CV20211 cited

ScanBank: A Benchmark Dataset for Figure Extraction from Scanned Electronic Theses and Dissertations

Sampanna Yashwant Kahu, William A. Ingram, Edward A. Fox +1

We focus on electronic theses and dissertations (ETDs), aiming to improve access and expand their utility, since more than 6 million are publicly available, and they constitute an…

cs.CL2021

Extractive Research Slide Generation Using Windowed Labeling Ranking

Athar Sefid, Jian Wu, Prasenjit Mitra +1

Presentation slides describing the content of scientific and technical papers are an efficient and effective way to present that work. However, manually generating presentation sli…

cs.DL2021

Predicting the Reproducibility of Social and Behavioral Science Papers Using Supervised Learning Models

Jian Wu, Rajal Nivargi, Sree Sai Teja Lanka +8

In recent years, significant effort has been invested verifying the reproducibility and robustness of research claims in social and behavioral sciences (SBS), much of which has inv…

cs.DL2019

Cleaning Noisy and Heterogeneous Metadata for Record Linking Across Scholarly Big Datasets

Athar Sefid, Jian Wu, Allen C. Ge +5

Automatically extracted metadata from scholarly documents in PDF formats is usually noisy and heterogeneous, often containing incomplete fields and erroneous values. One common way…