4 papers
Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training
Jingwei Zuo, Cong Zeng, Ilyas Chahed +6
The training paradigm of large language models has shifted from traditional one-pass training to multi-epoch training, as reasonable reuse of limited high-quality data can improve…
DEG: Efficient Hybrid Vector Search Using the Dynamic Edge Navigation Graph
Ziqi Yin, Jianyang Gao, Pasquale Balsebre +2
Bimodal data, such as image-text pairs, has become increasingly prevalent in the digital era. The Hybrid Vector Query (HVQ) is an effective approach for querying such data and has…
LAMP: A Language Model on the Map
Pasquale Balsebre, Weiming Huang, Gao Cong
Large Language Models (LLMs) are poised to play an increasingly important role in our lives, providing assistance across a wide array of tasks. In the geospatial domain, LLMs have…
City Foundation Models for Learning General Purpose Representations from OpenStreetMap
Pasquale Balsebre, Weiming Huang, Gao Cong +1
Pre-trained Foundation Models (PFMs) have ushered in a paradigm-shift in Artificial Intelligence, due to their ability to learn general-purpose representations that can be readily…