7 papers · 1 filter
Detecting Non-Membership in LLM Training Data via Rank Correlations
Pranav Shetty, Mirazul Haque, Zhiqiang Ma +1
As large language models (LLMs) are trained on increasingly vast and opaque text corpora, determining which data contributed to training has become essential for copyright enforcem…
ExStrucTiny: A Benchmark for Schema-Variable Structured Information Extraction from Document Images
Mathieu Sibue, Andres Muñoz Garza, Samuel Mensah +4
Enterprise documents, such as forms and reports, embed critical information for downstream applications like data archiving, automated workflows, and analytics. Although generalist…
Entropy-Gated Branching for Efficient Test-Time Reasoning
Xianzhi Li, Ethan Callanan, Abdellah Ghassel +1
Test-time compute methods can significantly improve the reasoning capabilities and problem-solving accuracy of large language models (LLMs). However, these approaches require subst…
Perturb Your Data: Paraphrase-Guided Training Data Watermarking
Pranav Shetty, Mirazul Haque, Petr Babkin +3
Training data detection is critical for enforcing copyright and data licensing, as Large Language Models (LLM) are trained on massive text corpora scraped from the internet. We pre…
The Oracle Has Spoken: A Multi-Aspect Evaluation of Dialogue in Pythia
Zixun Chen, Petr Babkin, Akshat Gupta +2
Dialogue is one of the landmark abilities of large language models (LLMs). Despite its ubiquity, few studies actually distinguish specific ingredients underpinning dialogue behavio…
CoCoLex: Confidence-guided Copy-based Decoding for Grounded Legal Text Generation
Santosh T. Y. S. S, Youssef Tarek Elkhayat, Oana Ichim +5
Due to their ability to process long and complex contexts, LLMs can offer key benefits to the Legal domain, but their adoption has been hindered by their tendency to generate unfai…