3 papers
cs.CL2026
Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
NVIDIA, :, Aaron Blakeman +571
We introduce Nemotron 3 Ultra, a 550 billion total and 55 billion active parameter Mixture-of-Experts Hybrid Mamba-Attention language model. We pre-trained Nemotron 3 Ultra on 20 t…
cs.CV2026
MMLongEmbed: Benchmarking Multimodal Embedding Models in Long-Context Scenarios
Haitian Wang, Ruoxi Sun, Quantong Qiu +5
Recent advancements have significantly expanded the theoretical context windows of Multimodal Embedding Models (MEMs). However, larger context windows do not necessarily translate…
cs.AI2026
Harnessing Pre-Resolution Signals for Future Prediction Agents
Chuyang Wei, Maohang Gao, Zhixin Han +12
Many high-stakes decisions depend on forecasts made before outcomes are known. In this future prediction setting, the central challenge is that public evidence evolves over time, w…