activity
20242026
collaborators
Showing 2024Show all

5 papers · 1 filter

cs.CL2024

RedPajama: an Open Dataset for Training Large Language Models

Maurice Weber, Daniel Fu, Quentin Anthony +16

Large language models are increasingly becoming a cornerstone technology in artificial intelligence, the sciences, and society as a whole, yet the optimal strategies for dataset co…

cs.LG2024

Spectra: Surprising Effectiveness of Pretraining Ternary Language Models at Scale

Ayush Kaushal, Tejas Vaidhya, Arnab Kumar Mondal +3

Rapid advancements in GPU computational power has outpaced memory capacity and bandwidth growth, creating bottlenecks in Large Language Model (LLM) inference. Post-training quantiz…

cs.LG2024

Simple and Scalable Strategies to Continually Pre-train Large Language Models

Adam Ibrahim, Benjamin Thérien, Kshitij Gupta +5

Large language models (LLMs) are routinely pre-trained on billions of tokens, only to start the process over again once new data becomes available. A much more efficient solution i…

cs.AI2024

Interpretability in Action: Exploratory Analysis of VPT, a Minecraft Agent

Karolis Jucys, George Adamopoulos, Mehrab Hamidi +7

Understanding the mechanisms behind decisions taken by large foundation models in sequential decision making tasks is critical to ensuring that such systems operate transparently a…

cs.CV2024

Towards Adversarially Robust Vision-Language Models: Insights from Design Choices and Prompt Formatting Techniques

Rishika Bhagwatkar, Shravan Nayak, Reza Bayat +4

Vision-Language Models (VLMs) have witnessed a surge in both research and real-world applications. However, as they are becoming increasingly prevalent, ensuring their robustness a…