Web Archives Metadata Generation with GPT-4o: Challenges and Insights
arXiv:2411.05409 · doi:10.5860/ital.v44i2.17305 10.5860/ital.v44i2.17305 10.5860/ital.v44i2.17305
Abstract
Current metadata creation for web archives is time consuming and costly due to reliance on human effort. This paper explores the use of gpt-4o for metadata generation within the Web Archive Singapore, focusing on scalability, efficiency, and cost effectiveness. We processed 112 Web ARChive (WARC) files using data reduction techniques, achieving a notable 99.9% reduction in metadata generation costs. By prompt engineering, we generated titles and abstracts, which were evaluated both intrinsically using Levenshtein Distance and BERTScore, and extrinsically with human cataloguers using McNemar's test. Results indicate that while our method offers significant cost savings and efficiency gains, human curated metadata maintains an edge in quality. The study identifies key challenges including content inaccuracies, hallucinations, and translation issues, suggesting that Large Language Models (LLMs) should serve as complements rather than replacements for human cataloguers. Future work will focus on refining prompts, improving content filtering, and addressing privacy concerns through experimentation with smaller models. This research advances the integration of LLMs in web archiving, offering valuable insights into their current capabilities and outlining directions for future enhancements. The code is available at https://github.com/masamune-prog/warc2summary for further development and use by institutions facing similar challenges.
Published in Information Technology and Libraries, Vol. 44, No. 2, June 2025
References in corpus (13)
- Learning Transferable Visual Models From Natural Language Supervision
- Training language models to follow instructions with human feedback
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Language Models are Few-Shot Learners
- BERTScore: Evaluating Text Generation with BERT
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- Graph of Thoughts: Solving Elaborate Problems with Large Language Models
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Extracting Training Data from Large Language Models
- PMC-LLaMA: Towards Building Open-source Language Models for Medicine
- FinGPT: Democratizing Internet-scale Data for Financial Large Language Models
- Web Archives Metadata Generation with GPT-4o: Challenges and Insights
- An Experiment with the Use of ChatGPT for LCSH Subject Assignment on Electronic Theses and Dissertations