activity
20242026
collaborators

6 papers

cs.IR2026

Utilizing Metadata for Better Retrieval-Augmented Generation

Raquib Bin Yousuf, Shengzhe Xu, Mandar Sharma +3

Retrieval-Augmented Generation systems depend on retrieving semantically relevant document chunks to support accurate, grounded outputs from large language models. In structured an…

cs.LG2025

Optimizing Product Provenance Verification using Data Valuation Methods

Raquib Bin Yousuf, Hoang Anh Just, Shengzhe Xu +8

Determining and verifying product provenance remains a critical challenge in global supply chains, particularly as geopolitical conflicts and shifting borders create new incentives…

cs.CL2025

Can an LLM Induce a Graph? Investigating Memory Drift and Context Length

Raquib Bin Yousuf, Aadyant Khatri, Shengzhe Xu +2

Recently proposed evaluation benchmarks aim to characterize the effective context length and the forgetting tendencies of large language models (LLMs). However, these benchmarks of…

cs.LG2025

Chasing the Timber Trail: Machine Learning to Reveal Harvest Location Misrepresentation

Shailik Sarkar, Raquib Bin Yousuf, Linhan Wang +9

Illegal logging poses a significant threat to global biodiversity, climate stability, and depresses international prices for legal wood harvesting and responsible forest products t…

cs.LG2025

Why LLMs Are Bad at Synthetic Table Generation (and what to do about it)

Shengzhe Xu, Cho-Ting Lee, Mandar Sharma +3

Synthetic data generation is integral to ML pipelines, e.g., to augment training data, replace sensitive information, and even to power advanced platforms like DeepSeek. While LLMs…

cs.CL2024

LLM Augmentations to support Analytical Reasoning over Multiple Documents

Raquib Bin Yousuf, Nicholas Defelice, Mandar Sharma +2

Building on their demonstrated ability to perform a variety of tasks, we investigate the application of large language models (LLMs) to enhance in-depth analytical reasoning within…