activity
20242026
collaborators

6 papers

cs.AI2026

Hack-Verifiable Terminal Bench: Evaluating Reward Hacking in Terminal Tasks

Amit Roth, Ivan Bercovich, Yonathan Efroni

As agents grow more capable and autonomous, their tendency to reward hack, satisfying a task's checks while violating its intent, becomes an increasingly important failure mode. Me…

cs.LG2026

Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale

Amit Roth, Ankur Samanta, Matan Halevy +2

Aligning autonomous agents with human intent remains a central challenge in modern AI. A key manifestation of this challenge is reward hacking, whereby agents appear successful und…

cs.CL2025

Scaling Analysis of Interleaved Speech-Text Language Models

Gallil Maimon, Michael Hassid, Amit Roth +1

Existing Speech Language Model (SLM) scaling analysis paints a bleak picture. It predicts that SLMs require much more compute and data compared to text, leading some to question th…

cs.SD2024

Salmon: A Suite for Acoustic Language Model Evaluation

Gallil Maimon, Amit Roth, Yossi Adi

Speech language models have recently demonstrated great potential as universal speech processing systems. Such models have the ability to model the rich acoustic information existi…

cs.CL2024

A Language Modeling Approach to Diacritic-Free Hebrew TTS

Amit Roth, Arnon Turetzky, Yossi Adi

We tackle the task of text-to-speech (TTS) in Hebrew. Traditional Hebrew contains Diacritics, which dictate the way individuals should pronounce given words, however, modern Hebrew…

cs.CL2024

HebDB: a Weakly Supervised Dataset for Hebrew Speech Processing

Arnon Turetzky, Or Tal, Yael Segal-Feldman +9

We present HebDB, a weakly supervised dataset for spoken language processing in the Hebrew language. HebDB offers roughly 2500 hours of natural and spontaneous speech recordings in…