Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
LSHBloom: Memory-efficient, Extreme-scale Document Deduplication
Arham Khan, Robert Underwood, Carlo Siebenschuh +7
Contemporary large language model (LLM) training pipelines require the assembly of internet-scale databases full of text data from a variety of sources (e.g., web, academic, and pu…
cs.LG2025
BioNeMo Framework: a modular, high-performance library for AI model development in drug discovery
Peter St. John, Dejun Lin, Polina Binder +89
Artificial Intelligence models encoding biology and chemistry are opening new routes to high-throughput and high-quality in-silico drug development. However, their training increas…