Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Abra: Scaling Diffusion Image Training
Kyle Chickering, Wei-An Lin, Swayam Bhanded +5
Compute-optimal scaling laws guide the training of frontier language models yet remain largely unexplored for visual generation. We present a systematic scaling law study for text-…
cs.LG2025
MINERVA: Evaluating Complex Video Reasoning
Arsha Nagrani, Sachit Menon, Ahmet Iscen +9
Multimodal LLMs are turning their focus to video benchmarks, however most video benchmarks only provide outcome supervision, with no intermediate or interpretable reasoning steps.…