3 papers
cs.DB2026
Filtered Vector Search in a Disaggregated Lakehouse: Composing Table-Format Pruning with Per-File ANN
Rakesh Jain, Thomas Griffin, Syed Zawad
Approximate nearest-neighbor (ANN) search increasingly runs alongside structured data - "find the 10 nearest documents where tenant='acme' AND lang='en'" - yet similarity and filte…
cs.CL2025
GneissWeb: Preparing High Quality Data for LLMs at Scale
Hajar Emami Gohari, Swanand Ravindra Kadhe, Syed Yousaf Shah +29
Data quantity and quality play a vital role in determining the performance of Large Language Models (LLMs). High-quality data, in particular, can significantly boost the LLM's abil…
cs.CV2025
Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence
Granite Vision Team, Leonid Karlinsky, Assaf Arbelle +60
We introduce Granite Vision, a lightweight large language model with vision capabilities, specifically designed to excel in enterprise use cases, particularly in visual document un…