5 papers
Toward Robust URL Extraction for Open Science: A Study of arXiv File Formats and Temporal Trends
Rochana R. Obadage, Lamia Salsabil, Sawood Alam +4
In this work, we study how URL extraction results depend on input format. We compiled a pilot dataset by extracting URLs from 10 arXiv papers and used the same heuristic method to…
When LLMs Disagree: Diagnosing Relevance Filtering Bias and Retrieval Divergence in SDG Search
William A. Ingram, Bipasha Banerjee, Edward A. Fox
Large language models (LLMs) are increasingly used to assign document relevance labels in information retrieval pipelines, especially in domains lacking human-labeled data. However…
VTechAGP: An Academic-to-General-Audience Text Paraphrase Dataset and Benchmark Models
Ming Cheng, Jiaying Gong, Chenhan Yuan +3
Existing text simplification or paraphrase datasets mainly focus on sentence-level text generation in a general domain. These datasets are typically developed without using domain…
Automating Chapter-Level Classification for Electronic Theses and Dissertations
Bipasha Banerjee, William A. Ingram, Edward A. Fox
Traditional archival practices for describing electronic theses and dissertations (ETDs) rely on broad, high-level metadata schemes that fail to capture the depth, complexity, and…
Agentic AI for Improving Precision in Identifying Contributions to Sustainable Development Goals
William A. Ingram, Bipasha Banerjee, Edward A. Fox
As research institutions increasingly commit to supporting the United Nations' Sustainable Development Goals (SDGs), there is a pressing need to accurately assess their research ou…