activity
20242026
collaborators

5 papers

cs.CL2026

Enabling Intrinsic Reasoning over Dense Geospatial Embeddings with DFR-Gemma

Xuechen Zhang, Aviv Slobodkin, Joydeep Paul +4

Representation learning for geospatial and spatio-temporal data plays a critical role in enabling general-purpose geospatial intelligence. Recent geospatial foundation models, such…

cs.IR2026

Utilizing Metadata for Better Retrieval-Augmented Generation

Raquib Bin Yousuf, Shengzhe Xu, Mandar Sharma +3

Retrieval-Augmented Generation systems depend on retrieving semantically relevant document chunks to support accurate, grounded outputs from large language models. In structured an…

cs.CL2025

Can an LLM Induce a Graph? Investigating Memory Drift and Context Length

Raquib Bin Yousuf, Aadyant Khatri, Shengzhe Xu +2

Recently proposed evaluation benchmarks aim to characterize the effective context length and the forgetting tendencies of large language models (LLMs). However, these benchmarks of…

cs.LG2025

Why LLMs Are Bad at Synthetic Table Generation (and what to do about it)

Shengzhe Xu, Cho-Ting Lee, Mandar Sharma +3

Synthetic data generation is integral to ML pipelines, e.g., to augment training data, replace sensitive information, and even to power advanced platforms like DeepSeek. While LLMs…

cs.CL2024

LLM Augmentations to support Analytical Reasoning over Multiple Documents

Raquib Bin Yousuf, Nicholas Defelice, Mandar Sharma +2

Building on their demonstrated ability to perform a variety of tasks, we investigate the application of large language models (LLMs) to enhance in-depth analytical reasoning within…