papers

Publications (5)

cs.CL2025

GneissWeb: Preparing High Quality Data for LLMs at Scale

Hajar Emami Gohari, Swanand Ravindra Kadhe, Syed Yousaf Shah +29

Data quantity and quality play a vital role in determining the performance of Large Language Models (LLMs). High-quality data, in particular, can significantly boost the LLM's abil…

cs.CR2022

Attack Techniques and Threat Identification for Vulnerabilities

Constantin Adam, Muhammed Fatih Bulut, Daby Sow +3

Modern organizations struggle with insurmountable number of vulnerabilities that are discovered and reported by their network and application vulnerability scanners. Therefore, pri…

cs.DC2026

FOLD: Fuzzy Online Deduplication for Very Large Evolving Datasets via Approximate Nearest Neighbor Search

Nelson Bore, Pritish Mishra, Constantin Adam +2

Fuzzy deduplication is key to constructing large language model training corpora. However, classic Locality-Sensitive Hashing (LSH) pipelines scale poorly as corpora grow and are i…

cs.CR2022

Partially Trusting the Service Mesh Control Plane

Constantin Adam, Abdulhamid Adebayo, Hubertus Franke +4

Zero Trust is a novel cybersecurity model that focuses on continually evaluating trust to prevent the initiation and horizontal spreading of attacks. A cloud-native Service Mesh is…

cs.AI2024

Data-Prep-Kit: getting your data ready for LLM application development

David Wood, Boris Lublinsky, Alexy Roytman +21

Data preparation is the first and a very important step towards any Large Language Model (LLM) development. This paper introduces an easy-to-use, extensible, and scale-flexible ope…