paper

Limits of Predictability in Civil Litigation

arXiv:2605.06151

Abstract

Legal practice routinely relies on informal assessments of case strength, yet no large-scale empirical benchmark exists for how predictable civil-litigation outcomes actually are. Civil litigation unfolds through sequential filings, and parties may settle at any stage, yet most computational studies of legal prediction observe cases only after resolution, leaving open whether outcomes are predictable beforehand. Using 102{,}721 U.S.\ civil cases and 835{,}190 court filings from 1996 to 2022, we model each case as it evolves, predicting plaintiff win, plaintiff loss, or settlement at each stage from structured, textual, and institutional features available up to that point. The classifier achieves class-specific AUC values of 0.74--0.81 and up to 97\% accuracy for high-confidence predictions, providing a large-scale benchmark for litigation predictability before resolution. We characterize heterogeneity in predictability using case complexity, defined as the entropy of the predicted outcome distribution. Complexity is systematically higher in cases involving corporate parties and in cases only weakly anchored to precedent. Richer information improves prediction mainly in low-complexity cases, with diminishing returns as complexity rises: some disputes are hard to predict not for lack of information, but because their outcomes are genuinely less determinate. Complexity also rises as litigation progresses, indicating that additional filings can sustain or amplify uncertainty rather than resolve it. Settlement rates follow an inverted U-shape in complexity, peaking at intermediate uncertainty and declining at both extremes. These findings suggest that predictive uncertainty is not mere model error, but a structured signal of legal complexity, litigation dynamics, and how disputes are resolved.

Limits of Predictability in Civil Litigation · wovepaper