collaborators

5 papers

cs.IR2025

Exploiting the Randomness of Large Language Models (LLM) in Text Classification Tasks: Locating Privileged Documents in Legal Matters

Keith Huffman, Jianping Zhang, Nathaniel Huber-Fliflet +2

In legal matters, text classification models are most often used to filter through large datasets in search of documents that meet certain pre-selected criteria like relevance to a…

cs.IR2025

Leveraging Machine Learning and Large Language Models for Automated Image Clustering and Description in Legal Discovery

Qiang Mao, Fusheng Wei, Robert Neary +4

The rapid increase in digital image creation and retention presents substantial challenges during legal discovery, digital archive, and content management. Corporations and legal t…

cs.IR2025

A Comparative Study of Retrieval Methods in Azure AI Search

Qiang Mao, Han Qin, Robert Neary +4

Increasingly, attorneys are interested in moving beyond keyword and semantic search to improve the efficiency of how they find key information during a document review task. Large…

cs.IR2025

Detecting Privileged Documents by Ranking Connected Network Entities

Jianping Zhang, Han Qin, Nathaniel Huber-Fliflet

This paper presents a link analysis approach for identifying privileged documents by constructing a network of human entities derived from email header metadata. Entities are class…

cs.IR2025

Empirical Evaluation of Embedding Models in the Context of Text Classification in Document Review in Construction Delay Disputes

Fusheng Wei, Robert Neary, Han Qin +2

Text embeddings are numerical representations of text data, where words, phrases, or entire documents are converted into vectors of real numbers. These embeddings capture semantic…