activity
20182026
most citedDocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis

139 citations · 193 across the 14 of their papers we have counts for

collaborators
Showing 2024Show all

5 papers · 1 filter

cs.IR2024★ 2 cited

Know Your RAG: Dataset Taxonomy and Generation Strategies for Evaluating RAG Systems

Rafael Teixeira de Lima, Shubham Gupta, Cesar Berrospi +4

Retrieval Augmented Generation (RAG) systems are a widespread application of Large Language Models (LLMs) in the industry. While many tools exist empowering developers to build the…

cs.AI2024

Data-Prep-Kit: getting your data ready for LLM application development

David Wood, Boris Lublinsky, Alexy Roytman +21

Data preparation is the first and a very important step towards any Large Language Model (LLM) development. This paper introduces an easy-to-use, extensible, and scale-flexible ope…

cs.CL2024★ 12 cited

Docling Technical Report

Christoph Auer, Maksym Lysak, Ahmed Nassar +16

This technical report introduces Docling, an easy to use, self-contained, MIT-licensed open-source package for PDF document conversion. It is powered by state-of-the-art specialize…

cs.CL2024★ 5 cited

Statements: Universal Information Extraction from Tables with Large Language Models for ESG KPIs

Lokesh Mishra, Sohayl Dhibi, Yusik Kim +4

Environment, Social, and Governance (ESG) KPIs assess an organization's performance on issues such as climate change, greenhouse gas emissions, water consumption, waste management,…

cs.CL2024★ 2 cited

INDUS: Effective and Efficient Language Models for Scientific Applications

Bishwaranjan Bhattacharjee, Aashka Trivedi, Masayasu Muraoka +33

Large language models (LLMs) trained on general domain corpora showed remarkable results on natural language processing (NLP) tasks. However, previous research demonstrated LLMs tr…