From the 1 of 4 linked papers with an AI index.
4 papers
Prompt-Induced Waste in Coding Agents: Reasoning, Effort, Harness Design, and End-to-End Cost
Sarel Weinberger, Amir Hozez
Coding-agent efficiency cannot be characterized by token count or model price alone. End-to-end cost and task success depend jointly on prompt semantics, inference effort, harness…
Token Reduction Is Not Cost Reduction
Sarel Weinberger, Amir Hozez
The paper studies whether context‑reduction techniques for API‑based coding agents actually lower the billed cost of using large language models, showing that token reduction often…
Mixture of Experts for Low-Resource LLMs
Ori Bar Joseph, Smadar Arvatz, Noam Kayzer +2
Mixture-of-Experts (MoE) architectures enable efficient model scaling, yet expert routing behavior across underrepresented languages remains poorly understood. We analyze routing d…
HEBATRON: A Hebrew-Specialized Open-Weight Mixture-of-Experts Language Model
Noam Kayzer, Dan Revital, Ori Bar Joseph +10
We present Hebatron, a Hebrew-specialized open-weight large language model built on the NVIDIA Nemotron-3 sparse Mixture-of-Experts architecture. Training employs a three-phase eas…