16 papers
S-EMBER: A Large-Scale Benchmark for Streaming Egocentric Memory Retrieval
Xiaodong Wang, Xuanyi Zhao, Pedro Rodriguez +7
As wearable devices enable continuous first-person recording, AI assistants must reason across long time horizons to recall past experiences-a capability known as episodic memory.…
The Llama 4 Herd: Architecture, Training, Evaluation, and Deployment Notes
Redacted by arXiv
This document consolidates publicly reported technical details about Metas Llama 4 model family. It summarizes (i) released variants (Scout and Maverick) and the broader herd conte…
In Pursuit of Pixel Supervision for Visual Pre-training
Lihe Yang, Shang-Wen Li, Yang Li +5
At the most basic level, pixels are the source of the visual information through which we perceive the world. Pixels contain information at all levels, ranging from low-level attri…
Demystifying CLIP Data
Hu Xu, Saining Xie, Xiaoqing Ellen Tan +7
Contrastive Language-Image Pre-training (CLIP) is an approach that has advanced research and applications in computer vision, fueling modern recognition systems and generative mode…
Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
Feiyang Kang, Newsha Ardalani, Michael Kuchnik +7
Training data plays a crucial role in Large Language Models (LLM) scaling, yet high quality data is of limited supply. Synthetic data techniques offer a potential path toward sides…
DepthLM: Metric Depth From Vision Language Models
Zhipeng Cai, Ching-Feng Yeh, Hu Xu +7
Vision language models (VLMs) can flexibly address various vision tasks through text interactions. Although successful in semantic understanding, state-of-the-art VLMs including GP…