activity
20182026
most citedBeyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

565 citations · 717 across the 41 of their papers we have counts for

collaborators

61 papers

cs.CL2026

CT: Counterfactual Causal Reasoning for Sentiment Shifts in Social-Media Conversation Trees

S M Rafiuddin, Atriya Sen

Sentiment in social-media threads does not only vary across posts; it shifts as users react to claims, corrections, evidence, and hostility within a branching reply tree. We study…

cs.CV2026

Modality Maturity Index: A benchmark for assessing multimodal capabilities of omni models

Rohit Patel, Dieuwke Hupkes, Sloan Strader

Frontier language models are increasingly marketed as omni systems that can perceive and respond across modalities. Existing evaluation frameworks, however, focus almost exclusivel…

cs.LG2026

LPDS: Evaluating LLM Robustness Through Logic-Preserving Difficulty Scaling

Philipp Mondorf, Samuel J. Bell, Jesse Dodge +1

As large language models (LLMs) are increasingly deployed to perform tasks with minimal human oversight, it is crucial that these models operate robustly. In particular, a model th…

cs.CL2026

Gender Bias in MT for a Genderless Language: New Benchmarks for Basque

Amaia Murillo, Olatz-Perez-de-Viñaspre, Naiara Perez

Large language models (LLMs) and machine translation (MT) systems are increasingly used in our daily lives, but their outputs can reproduce gender bias present in the training data…

cs.SE2026★ 1 cited

The Llama 4 Herd: Architecture, Training, Evaluation, and Deployment Notes

Redacted by arXiv

This document consolidates publicly reported technical details about Metas Llama 4 model family. It summarizes (i) released variants (Scout and Maverick) and the broader herd conte…

cs.CL2025

From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation

Seokhee Hong, Sunkyoung Kim, Guijin Son +3

The development of Large Language Models (LLMs) requires robust benchmarks that encompass not only academic domains but also industrial fields to effectively evaluate their applica…