3 papers
cs.LG2026
SHRED: Retain-Set-Free Unlearning via Self-Distillation with Logit Demotion
Zizhao Hu, Ameya Godbole, Johnny Tian-Zheng Wei +3
Machine unlearning for large language models (LLMs) aims to selectively remove memorized content such as private data, copyrighted text, or hazardous knowledge, without costly full…
cs.CL2026
Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing
Minglai Yang, Xinyan Velocity Yu, Pengyuan Li +22
Document parsing and recognition are fundamental capabilities for vision-language models (VLMs) and document processing systems. However, existing Optical Character Recognition (OC…
cs.CL2025
Phonological Representation Learning for Isolated Signs Improves Out-of-Vocabulary Generalization
Lee Kezar, Zed Sehyr, Jesse Thomason
Sign language datasets are often not representative in terms of vocabulary, underscoring the need for models that generalize to unseen signs. Vector quantization is a promising app…