most citedScaling Knowledge Graph Construction through Synthetic Data Generation and Distillation

1 citations · 1 across the 5 of their papers we have counts for

collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL2026

Dont Stop Early: Scalable Enterprise Deep Research with Controlled Information Flow and Evidence-Aware Termination

Prafulla Kumar Choubey, Kung-Hsiang Huang, Pranav Narayanan Venkit +5

Enterprise deep research often fails to produce decision-ready reports due to uneven information coverage, context explosion, and premature stopping. We propose a scalable Enterpri…

cs.CL20261 cited

Scaling Knowledge Graph Construction through Synthetic Data Generation and Distillation

Prafulla Kumar Choubey, Xin Su, Man Luo +9

Document-level knowledge graph (KG) construction faces a fundamental scaling challenge: existing methods either rely on expensive large language models (LLMs), making them economic…

cs.CL2026

UNIDOC-BENCH: A Unified Benchmark for Document-Centric Multimodal RAG

Xiangyu Peng, Can Qin, Zeyuan Chen +3

Multimodal retrieval-augmented Generation (MM-RAG) is a key approach for applying large language models (LLMs) and agents to real-world knowledge bases, yet current evaluations are…

cs.CL2025

Benchmarking Deep Search over Heterogeneous Enterprise Data

Prafulla Kumar Choubey, Xiangyu Peng, Shilpa Bhagavath +3

We present a new benchmark for evaluating Deep Search--a realistic and complex form of retrieval-augmented generation (RAG) that requires source-aware, multi-hop reasoning over div…

cs.CL2025

Unanswerability Evaluation for Retrieval Augmented Generation

Xiangyu Peng, Prafulla Kumar Choubey, Caiming Xiong +1

Existing evaluation frameworks for retrieval-augmented generation (RAG) systems focus on answerable queries, but they overlook the importance of appropriately rejecting unanswerabl…

cs.CL2025

ReGenesis: LLMs can Grow into Reasoning Generalists via Self-Improvement

Xiangyu Peng, Congying Xia, Xinyi Yang +3

Post-training Large Language Models (LLMs) with explicit reasoning trajectories can enhance their reasoning abilities. However, acquiring such high-quality trajectory data typicall…