papers

Publications (84)

cs.CL2023

A Survey on Multi-modal Summarization

Anubhav Jangra, Sourajit Mukherjee, Adam Jatowt +2

The new era of technology has brought us to the point where it is convenient for people to share their opinions over an abundance of platforms. These platforms have a provision for…

cs.CL2025

Evaluating Answer Reranking Strategies in Time-sensitive Question Answering

Mehmet Kardan, Bhawna Piryani, Adam Jatowt

Despite advancements in state-of-the-art models and information retrieval techniques, current systems still struggle to handle temporal information and to correctly answer detailed…

cs.CL2023

Archive TimeLine Summarization (ATLS): Conceptual Framework for Timeline Generation over Historical Document Collections

Nicolas Gutehrlé, Antoine Doucet, Adam Jatowt

Archive collections are nowadays mostly available through search engines interfaces, which allow a user to retrieve documents by issuing queries. The study of these collections may…

cs.IR2026

Argus-Retriever: Vision-LLM Late-Interaction Retrieval with Region-Aware Query-Conditioned MoE for Visual Document Retrieval

Abdelrahman Abdallah, Mahmoud Abdalla, Mohammed Ali +1

Late-interaction vision-language retrievers represent each document page as many visual token embeddings and score queries with MaxSim. In systems such as ColPali, ColQwen, ColNomi…

cs.CL2025

Detecting Future-related Contexts of Entity Mentions

Puneet Prashar, Krishna Mohan Shukla, Adam Jatowt

The ability to automatically identify whether an entity is referenced in a future context can have multiple applications including decision making, planning and trend forecasting.…

cs.CL2025

Analyzing the Role of Context in Forecasting with Large Language Models

Gerrit Mutschlechner, Adam Jatowt

This study evaluates the forecasting performance of recent language models (LLMs) on binary forecasting questions. We first introduce a novel dataset of over 600 binary forecasting…

cs.IR2026

Difficulty-Gated Fusion of Reasoning Views for Temporal Retrieval

Jamie Holdcroft, Abdelrahman Abdallah, Adam Jatowt

Reasoning-intensive temporal retrieval requires matching a query to documents whose relevance depends on shared temporal reasoning rather than lexical overlap. Expanding a query in…

cs.CV2025

ReceiptSense: Beyond Traditional OCR -- A Dataset for Receipt Understanding

Abdelrahman Abdallah, Mohamed Mounis, Mahmoud Abdalla +6

Multilingual OCR and information extraction from receipts remains challenging, particularly for complex scripts like Arabic. We introduce \dataset, a comprehensive dataset designed…

q-fin.GN2022

Predicting Companies' ESG Ratings from News Articles Using Multivariate Timeseries Analysis

Tanja Aue, Adam Jatowt, Michael Färber

Environmental, social and governance (ESG) engagement of companies moved into the focus of public attention over recent years. With the requirements of compulsory reporting being i…

cs.IR2025

Rankify: A Comprehensive Python Toolkit for Retrieval, Re-Ranking, and Retrieval-Augmented Generation

Abdelrahman Abdallah, Bhawna Piryani, Jamshid Mozafari +2

Retrieval, re-ranking, and retrieval-augmented generation (RAG) are critical components of modern applications in information retrieval, question answering, or knowledge-based text…

cs.CL2025

Automated Analysis of Sustainability Reports: Using Large Language Models for the Extraction and Prediction of EU Taxonomy-Compliant KPIs

Jonathan Schmoll, Adam Jatowt

The manual, resource-intensive process of complying with the EU Taxonomy presents a significant challenge for companies. While Large Language Models (LLMs) offer a path to automati…

cs.IR2025

RankArena: A Unified Platform for Evaluating Retrieval, Reranking and RAG with Human and LLM Feedback

Abdelrahman Abdallah, Mahmoud Abdalla, Bhawna Piryani +3

Evaluating the quality of retrieval-augmented generation (RAG) and document reranking systems remains challenging due to the lack of scalable, user-centric, and multi-perspective e…

cs.IR2026

BracketRank: Large Language Model Document Ranking via Reasoning-based Competitive Elimination

Abdelrahman Abdallah, Mohammed Ali, Bhawna Piryani +1

Reasoning-intensive retrieval requires deep semantic inference beyond surface-level keyword matching, posing a challenge for current LLM-based rerankers limited by context constrai…

cs.CL2024

Navigating the Landscape of Hint Generation Research: From the Past to the Future

Anubhav Jangra, Jamshid Mozafari, Adam Jatowt +1

Digital education has gained popularity in the last decade, especially after the COVID-19 pandemic. With the improving capabilities of large language models to reason and communica…

cs.LG2020

Joint Event Extraction along Shortest Dependency Paths using Graph Convolutional Networks

Ali Balali, Masoud Asadpour, Ricardo Campos +1

Event extraction (EE) is one of the core information extraction tasks, whose purpose is to automatically identify and extract information about incidents and their actors from text…

cs.CL2026

Event-Centric Human Value Understanding in News-Domain Texts: An Actor-Conditioned, Multi-Granularity Benchmark

Yao Wang, Xin Liu, Zhuochen Liu +5

Existing human value datasets do not directly support value understanding in factual news: many are actor-agnostic, rely on isolated utterances or synthetic scenarios, and lack exp…

cs.CL2026

Context Convergence Improves Answering Inferential Questions

Jamshid Mozafari, Bhawna Piryani, Adam Jatowt

While Large Language Models (LLMs) are widely used in open-domain Question Answering (QA), their ability to handle inferential questions-where answers must be derived rather than d…

cs.CL2023

Exploring the State of the Art in Legal QA Systems

Abdelrahman Abdallah, Bhawna Piryani, Adam Jatowt

Answering questions related to the legal domain is a complex task, primarily due to the intricate nature and diverse range of legal document systems. Providing an accurate answer t…

cs.CL2022

ArchivalQA: A Large-scale Benchmark Dataset for Open Domain Question Answering over Historical News Collections

Jiexin Wang, Adam Jatowt, Masatoshi Yoshikawa

In the last few years, open-domain question answering (ODQA) has advanced rapidly due to the development of deep learning techniques and the availability of large-scale QA datasets…

cs.CL2024

Generator-Retriever-Generator Approach for Open-Domain Question Answering

Abdelrahman Abdallah, Adam Jatowt

Open-domain question answering (QA) tasks usually require the retrieval of relevant information from a large corpus to generate accurate answers. We propose a novel approach called…

cs.IR2026

TEMPO: A Realistic Multi-Domain Benchmark for Temporal Reasoning-Intensive Retrieval

Abdelrahman Abdallah, Mohammed Ali, Muhammad Abdul-Mageed +1

Existing temporal QA benchmarks focus on simple fact-seeking queries from news corpora, while reasoning-intensive retrieval benchmarks lack temporal grounding. However, real-world…

cs.IR2026

MM-BRIGHT: A Multi-Task Multimodal Benchmark for Reasoning-Intensive Retrieval

Abdelrahman Abdallah, Mohamed Darwish Mounis, Mahmoud Abdalla +6

Existing retrieval benchmarks primarily consist of text-based queries where keyword or semantic matching is usually sufficient. Many real-world queries contain multimodal elements,…

cs.CL2026

Pretraining Exposure Explains Popularity Judgments in Large Language Models

Jamshid Mozafari, Bhawna Piryani, Adam Jatowt

Large language models (LLMs) exhibit systematic preferences for well-known entities, a phenomenon often attributed to popularity bias. However, the extent to which these preference…

cs.CL2026

Inferential Question Answering

Jamshid Mozafari, Hamed Zamani, Guido Zuccon +1

Despite extensive research on a wide range of question answering (QA) systems, most existing work focuses on answer containment-i.e., assuming that answers can be directly extracte…

cs.CL2025

WikiHint: A Human-Annotated Dataset for Hint Ranking and Generation

Jamshid Mozafari, Florian Gerhold, Adam Jatowt

The use of Large Language Models (LLMs) has increased significantly with users frequently asking questions to chatbots. In the time when information is readily accessible, it is cr…

cs.CL2024

Listening to Patients: A Framework of Detecting and Mitigating Patient Misreport for Medical Dialogue Generation

Lang Qin, Yao Zhang, Hongru Liang +2

Medical Dialogue Systems aim to provide automated healthcare support through patient-agent conversations. Previous efforts typically regard patients as ideal users -- one who accur…

cs.CL2024

Temporal Blind Spots in Large Language Models

Jonas Wallat, Adam Jatowt, Avishek Anand

Large language models (LLMs) have recently gained significant attention due to their unparalleled ability to perform various natural language processing tasks. These models, benefi…

cs.CL2025

Wrong Answers Can Also Be Useful: PlausibleQA -- A Large-Scale QA Dataset with Answer Plausibility Scores

Jamshid Mozafari, Abdelrahman Abdallah, Bhawna Piryani +1

Large Language Models (LLMs) are revolutionizing information retrieval, with chatbots becoming an important source for answering user queries. As by their design, LLMs prioritize g…

cs.CL2026

From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations

Pulkit Bansal, Raghvendra Kumar, Shakti Singh +2

In an era of rampant misinformation, generating reliable news explanations is vital, especially for under-represented languages like Hindi. Lacking robust automated tools, Hindi fa…

cs.CL2026

How often do Answers Change? Estimating Recency Requirements in Question Answering

Bhawna Piryani, Zehra Mert, Adam Jatowt

Large language models (LLMs) often rely on outdated knowledge when answering time-sensitive questions, leading to confident yet incorrect responses. Without explicit signals indica…

cs.CL2025

Evaluating Robustness of LLMs in Question Answering on Multilingual Noisy OCR Data

Bhawna Piryani, Jamshid Mozafari, Abdelrahman Abdallah +2

Optical Character Recognition (OCR) plays a crucial role in digitizing historical and multilingual documents, yet OCR errors - imperfect extraction of text, including character ins…

cs.DL2021

Change Summarization of Diachronic Scholarly Paper Collections by Semantic Evolution Analysis

Naman Paharia, Muhammad Syafiq Mohd Pozi, Adam Jatowt

The amount of scholarly data has been increasing dramatically over the last years. For newcomers to a particular science domain (e.g., IR, physics, NLP) it is often difficult to sp…

cs.IR2020

Citation Recommendation: Approaches and Datasets

Michael Färber, Adam Jatowt

Citation recommendation describes the task of recommending citations for a given text. Due to the overload of published scientific works in recent years on the one hand, and the ne…

cs.IR2026

Are LLM-Based Retrievers Worth Their Cost? An Empirical Study of Efficiency, Robustness, and Reasoning Overhead

Abdelrahman Abdallah, Jamie Holdcroft, Mohammed Ali +1

Large language model retrievers improve performance on complex queries, but their practical value depends on efficiency, robustness, and reliable confidence signals in addition to…

cs.LG2025

Wisdom of the Crowds in Forecasting: Forecast Summarization for Supporting Future Event Prediction

Anisha Saha, Adam Jatowt

Future Event Prediction (FEP) is an essential activity whose demand and application range across multiple domains. While traditional methods like simulations, predictive and time-s…

cs.CL2024

Transformers and Language Models in Form Understanding: A Comprehensive Review of Scanned Document Analysis

Abdelrahman Abdallah, Daniel Eberharter, Zoe Pfister +1

This paper presents a comprehensive survey of research works on the topic of form understanding in the context of scanned documents. We delve into recent advancements and breakthro…

cs.CL2025

A Study into Investigating Temporal Robustness of LLMs

Jonas Wallat, Abdelrahman Abdallah, Adam Jatowt +1

Large Language Models (LLMs) encapsulate a surprising amount of factual world knowledge. However, their performance on temporal questions and historical knowledge is limited becaus…

cs.CL2024

TriviaHG: A Dataset for Automatic Hint Generation from Factoid Questions

Jamshid Mozafari, Anubhav Jangra, Adam Jatowt

Nowadays, individuals tend to engage in dialogues with Large Language Models, seeking answers to their questions. In times when such answers are readily accessible to anyone, the s…

cs.AI2021

Generalized Relation Learning with Semantic Correlation Awareness for Link Prediction

Yao Zhang, Xu Zhang, Jun Wang +5

Developing link prediction models to automatically complete knowledge graphs has recently been the focus of significant research interest. The current methods for the link predicti…

cs.CL2019

Survey of Computational Approaches to Lexical Semantic Change

Nina Tahmasebi, Lars Borin, Adam Jatowt

Our languages are in constant flux driven by external factors such as cultural, societal and technological changes, as well as by only partially understood internal motivations. Wo…

cs.CL2025

ComplexTempQA:A 100m Dataset for Complex Temporal Question Answering

Raphael Gruber, Abdelrahman Abdallah, Michael Färber +1

We introduce \textsc{ComplexTempQA},\footnote{Dataset and code available at: https://github.com/DataScienceUIBK/ComplexTempQA} a large-scale dataset consisting of over 100 million…

cs.CL2021

A Neural Conversation Generation Model via Equivalent Shared Memory Investigation

Changzhen Ji, Yating Zhang, Xiaozhong Liu +4

Conversation generation as a challenging task in Natural Language Generation (NLG) has been increasingly attracting attention over the last years. A number of recent works adopted…

cs.IR2026

Negative Sampling Techniques in Information Retrieval: A Survey

Laurin Wischounig, Abdelrahman Abdallah, Adam Jatowt

Information Retrieval (IR) is fundamental to many modern NLP applications. The rise of dense retrieval (DR), using neural networks to learn semantic vector representations, has sig…

cs.MA2026

REGREACT: Self-Correcting Multi-Agent Pipelines for Structured Regulatory Information Extraction

Mohammed Ali, Abdelrahman Abdallah, Adam Jatowt

Extracting structured, machine-readable compliance criteria from regulatory documents remains an open challenge. Single-pass language models hallucinate structural elements, lose h…

cs.CL2024

Multi-hop Question Answering

Vaibhav Mavi, Anubhav Jangra, Adam Jatowt

The task of Question Answering (QA) has attracted significant research interest for long. Its relevance to language understanding and knowledge retrieval tasks, along with the simp…

cs.IR2020

ECIR 2020 Workshops: Assessing the Impact of Going Online

Sérgio Nunes, Suzanne Little, Sumit Bhatia +11

ECIR 2020 https://ecir2020.org/ was one of the many conferences affected by the COVID-19 pandemic. The Conference Chairs decided to keep the initially planned dates (April 14-17, 2…

cs.IR2026

RECOR: Reasoning-focused Multi-turn Conversational Retrieval Benchmark

Mohammed Ali, Abdelrahman Abdallah, Amit Agarwal +2

Existing benchmarks treat multi-turn conversation and reasoning-intensive retrieval separately, yet real-world information seeking requires both. To bridge this gap, we present a b…

cs.CL2024

AMuRD: Annotated Arabic-English Receipt Dataset for Key Information Extraction and Classification

Abdelrahman Abdallah, Mahmoud Abdalla, Mohamed Elkasaby +2

The extraction of key information from receipts is a complex task that involves the recognition and extraction of text from scanned receipts. This process is crucial as it enables…

cs.CL2025

Exploring NLP Benchmarks in an Extremely Low-Resource Setting

Ulin Nuha, Adam Jatowt

The effectiveness of Large Language Models (LLMs) diminishes for extremely low-resource languages, such as indigenous languages, primarily due to the lack of labeled data. Despite…

cs.SE2024

Is Your AI-Generated Code Really Safe? Evaluating Large Language Models on Secure Code Generation with CodeSecEval

Jiexin Wang, Xitong Luo, Liuwen Cao +5

Large language models (LLMs) have brought significant advancements to code generation and code repair, benefiting both novice and experienced developers. However, their training us…

cs.CL2024

ChroniclingAmericaQA: A Large-scale Question Answering Dataset based on Historical American Newspaper Pages

Bhawna Piryani, Jamshid Mozafari, Adam Jatowt

Question answering (QA) and Machine Reading Comprehension (MRC) tasks have significantly advanced in recent years due to the rapid development of deep learning techniques and, more…

cs.CL2026

Question Difficulty Estimation for Large Language Models via Answer Plausibility Scoring

Jamshid Mozafari, Bhawna Piryani, Adam Jatowt

Estimating question difficulty is a critical component in evaluating and improving large language models (LLMs) for question answering (QA). Existing approaches often rely on reada…

cs.CL2025

Navigating Tomorrow: Reliably Assessing Large Language Models Performance on Future Event Prediction

Petraq Nako, Adam Jatowt

Predicting future events is an important activity with applications across multiple fields and domains. For example, the capacity to foresee stock market trends, natural disasters,…

cs.CL2025

HintEval: A Comprehensive Framework for Hint Generation and Evaluation for Questions

Jamshid Mozafari, Bhawna Piryani, Abdelrahman Abdallah +1

Large Language Models (LLMs) are transforming how people find information, and many users turn nowadays to chatbots to obtain answers to their questions. Despite the instant access…

cs.IR2026

EXCISE: Query-Side Exclusion for Late-Interaction Retrieval

Mohammed Ali, Abdelrahman Abdallah, Adam Jatowt

Late-interaction retrievers handle exclusion queries poorly. When a user asks for X but not Z, the additive MaxSim score promotes documents covering Z, a problem we call exclusion…

cs.IR2025

The Impact of International Collaborations with Highly Publishing Countries in Computer Science

Alberto Gomez Espes, Michael Faerber, Adam Jatowt

This paper analyzes international collaborations in Computer Science, focusing on three major players: China, the European Union, and the United States. Drawing from a comprehensiv…

cs.AI2021

GMH: A General Multi-hop Reasoning Model for KG Completion

Yao Zhang, Hongru Liang, Adam Jatowt +4

Knowledge graphs are essential for numerous downstream natural language processing applications, but are typically incomplete with many facts missing. This results in research effo…

cs.CL2025

DeAR: Dual-Stage Document Reranking with Reasoning Agents via LLM Distillation

Abdelrahman Abdallah, Jamshid Mozafari, Bhawna Piryani +1

Large Language Models (LLMs) have transformed listwise document reranking by enabling global reasoning over candidate sets, yet single models often struggle to balance fine-grained…

cs.AI2026

MoCA-Agent: A Market-of-Claims Code Agent for Financial and Numerical Reasoning

Abdelrahman Abdallah, AbdelRahim A. Elmadany, Sameh Al Natour +3

Financial and tabular question answering requires more than fluent reasoning: answers must be grounded in the exact facts, formulas, units, signs, and scales that support them. A s…

cs.AI2023

An Overview Of Temporal Commonsense Reasoning and Acquisition

Georg Wenzel, Adam Jatowt

Temporal commonsense reasoning refers to the ability to understand the typical temporal context of phrases, actions, and events, and use it to reason over problems requiring such k…

cs.CL2025

How Good are LLM-based Rerankers? An Empirical Analysis of State-of-the-Art Reranking Models

Abdelrahman Abdallah, Bhawna Piryani, Jamshid Mozafari +2

In this work, we present a systematic and comprehensive empirical evaluation of state-of-the-art reranking methods, encompassing large language model (LLM)-based, lightweight conte…

cs.CV2025

Guess the Age of Photos: An Interactive Web Platform for Historical Image Age Estimation

Hasan Yucedag, Adam Jatowt

This paper introduces Guess the Age of Photos, a web platform engaging users in estimating the years of historical photographs through two gamified modes: Guess the Year (predictin…

cs.CL2025

Towards Effective Time-Aware Language Representation: Exploring Enhanced Temporal Understanding in Language Models

Jiexin Wang, Adam Jatowt, Yi Cai

In the evolving field of Natural Language Processing (NLP), understanding the temporal context of text is increasingly critical for applications requiring advanced temporal reasoni…

cs.CL2025

ASRank: Zero-Shot Re-Ranking with Answer Scent for Document Retrieval

Abdelrahman Abdallah, Jamshid Mozafari, Bhawna Piryani +1

Retrieval-Augmented Generation (RAG) models have drawn considerable attention in modern open-domain question answering. The effectiveness of RAG depends on the quality of the top r…

cs.SE2023

Enhancing Large Language Models for Secure Code Generation: A Dataset-driven Study on Vulnerability Mitigation

Jiexin Wang, Liuwen Cao, Xitong Luo +4

Large language models (LLMs) have brought significant advancements to code generation, benefiting both novice and experienced developers. However, their training using unsanitized…

cs.CL2023

BiTimeBERT: Extending Pre-Trained Language Representations with Bi-Temporal Information

Jiexin Wang, Adam Jatowt, Masatoshi Yoshikawa +1

Time is an important aspect of documents and is used in a range of NLP and IR tasks. In this work, we investigate methods for incorporating temporal information during pre-training…

cs.IR2025

SustainableQA: A Comprehensive Question Answering Dataset for Corporate Sustainability and EU Taxonomy Reporting

Mohammed Ali, Abdelrahman Abdallah, Adam Jatowt

The growing demand for corporate sustainability transparency, particularly under new regulations like the EU Taxonomy, necessitates precise data extraction from large, unstructured…

cs.CL2025

From Retrieval to Generation: Comparing Different Approaches

Abdelrahman Abdallah, Jamshid Mozafari, Bhawna Piryani +2

Knowledge-intensive tasks, particularly open-domain question answering (ODQA), document reranking, and retrieval-augmented language modeling, require a balance between retrieval ac…

cs.CL2024

Detecting Temporal Ambiguity in Questions

Bhawna Piryani, Abdelrahman Abdallah, Jamshid Mozafari +1

Detecting and answering ambiguous questions has been a challenging task in open-domain question answering. Ambiguous questions have different answers depending on their interpretat…

cs.CL2024

DynRank: Improving Passage Retrieval with Dynamic Zero-Shot Prompting Based on Question Classification

Abdelrahman Abdallah, Jamshid Mozafari, Bhawna Piryani +2

This paper presents DynRank, a novel framework for enhancing passage retrieval in open-domain question-answering systems through dynamic zero-shot question classification. Traditio…

cs.CL2026

It's High Time: A Survey of Temporal Question Answering

Bhawna Piryani, Abdelrahman Abdallah, Jamshid Mozafari +2

Time plays a critical role in how information is generated, retrieved, and interpreted. In this survey, we provide a comprehensive overview of Temporal Question Answering (TQA), a…

cs.CL2026

PARSE: An Open-Domain Reasoning Question Answering Benchmark for Persian

Jamshid Mozafari, Seyed Parsa Mousavinasab, Adam Jatowt

Reasoning-focused Question Answering (QA) has advanced rapidly with Large Language Models (LLMs), yet high-quality benchmarks for low-resource languages remain scarce. Persian, spo…

cs.CL2023

Measuring Variety, Balance, and Disparity: An Analysis of Media Coverage of the 2021 German Federal Election

Michael Färber, Jannik Schwade, Adam Jatowt

Determining and measuring diversity in news articles is important for a number of reasons, including preventing filter bubbles and fueling public discourse, especially before elect…

cs.CL2025

Evaluating List Construction and Temporal Understanding capabilities of Large Language Models

Alexandru Dumitru, V Venktesh, Adam Jatowt +1

Large Language Models (LLMs) have demonstrated immense advances in a wide range of natural language tasks. However, these models are susceptible to hallucinations and errors on par…

cs.CL2025

ClaimGen-CN: A Large-scale Chinese Dataset for Legal Claim Generation

Siying Zhou, Yiquan Wu, Hui Chen +6

Legal claims refer to the plaintiff's demands in a case and are essential to guiding judicial reasoning and case resolution. While many works have focused on improving the efficien…

cs.IR2020

Multi-Modal Summary Generation using Multi-Objective Optimization

Anubhav Jangra, Sriparna Saha, Adam Jatowt +1

Significant development of communication technology over the past few years has motivated research in multi-modal summarization techniques. A majority of the previous works on mult…

cs.IR2025

TempRetriever: Fusion-based Temporal Dense Passage Retrieval for Time-Sensitive Questions

Abdelrahman Abdallah, Bhawna Piryani, Jonas Wallat +2

Temporal awareness is crucial in many information retrieval tasks, particularly in scenarios where the relevance of documents depends on their alignment with the query's temporal c…

cs.DL2017

Towards Understanding the Evolution of the WWW Conference

Pavel Savov, Adam Jatowt, Radoslaw Nielek

The World Wide Web conference is a well-established and mature venue with an already long history. Over the years it has been attracting papers reporting many important research ac…

cs.IR2024

LLMTemporalComparator: A Tool for Analysing Differences in Temporal Adaptations of Large Language Models

Reinhard Friedrich Fritsch, Adam Jatowt

This study addresses the challenges of analyzing temporal discrepancies in large language models (LLMs) trained on data from different time periods. To facilitate the automatic exp…

cs.CL2024

Exploring Hint Generation Approaches in Open-Domain Question Answering

Jamshid Mozafari, Abdelrahman Abdallah, Bhawna Piryani +1

Automatic Question Answering (QA) systems rely on contextual information to provide accurate answers. Commonly, contexts are prepared through either retrieval-based or generation-b…

cs.CL2024

Temporal Validity Change Prediction

Georg Wenzel, Adam Jatowt

Temporal validity is an important property of text that is useful for many downstream applications, such as recommender systems, conversational AI, or story understanding. Existing…

cs.AI2022

Fact-Tree Reasoning for N-ary Question Answering over Knowledge Graphs

Yao Zhang, Peiyao Li, Hongru Liang +2

In the question answering(QA) task, multi-hop reasoning framework has been extensively studied in recent years to perform more efficient and interpretable answer reasoning on the K…

cs.CL2022

A Survey on Medical Document Summarization

Raghav Jain, Anubhav Jangra, Sriparna Saha +1

The internet has had a dramatic effect on the healthcare industry, allowing documents to be saved, shared, and managed digitally. This has made it easier to locate and share import…

cs.CL2024

ArabicaQA: A Comprehensive Dataset for Arabic Question Answering

Abdelrahman Abdallah, Mahmoud Kasem, Mahmoud Abdalla +4

In this paper, we address the significant gap in Arabic natural language processing (NLP) resources by introducing ArabicaQA, the first large-scale dataset for machine reading comp…