1 citations · 1 across the 6 of their papers we have counts for
7 papers
LLaTTE: Scaling Laws for Multi-Stage Sequence Modeling in Large-Scale Ads Recommendation
Lee Xiong, Zhirong Chen, Rahul Mayuranath +17
We present LLaTTE (LLM-Style Latent Transformers for Temporal Events), a scalable transformer architecture for production ads recommendation. Through systematic experiments, we dem…
Reasoning Models Ace the CFA Exams
Jaisal Patel, Yunzhe Chen, Kaiwen He +4
Previous research has reported that large language models (LLMs) demonstrate poor performance on the Chartered Financial Analyst (CFA) exams. However, recent reasoning models have…
GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
Tejal Patwardhan, Rachel Dias, Elizabeth Proehl +16
We introduce GDPval, a benchmark evaluating AI model capabilities on real-world economically valuable tasks. GDPval covers the majority of U.S. Bureau of Labor Statistics Work Acti…
Energy Consumption in Parallel Neural Network Training
Philipp Huber, David Li, Juan Pedro Gutiérrez Hermosillo Muriedas +4
The increasing demand for computational resources of training neural networks leads to a concerning growth in energy consumption. While parallelization has enabled upscaling model…
Modeling Urban Food Insecurity with Google Street View Images
David Li
Food insecurity is a significant social and public health issue that plagues many urban metropolitan areas around the world. Existing approaches to identifying food insecurity rely…
Large Language Model Compression via the Nested Activation-Aware Decomposition
Jun Lu, Tianyi Xu, Bill Ding +2
In this paper, we tackle the critical challenge of compressing large language models (LLMs) to facilitate their practical deployment and broader adoption. We introduce a novel post…