2 papers
cs.LG2026
StepShield: When, Not Whether to Intervene on Rogue Agents
Gloria Felicia, Zitha Sasindran, Jinfeng He +3
Agent safety benchmarks measure whether a monitor detects harm, not when. Yet timing is the difference between intervention and autopsy. We introduce StepShield, the first benchmar…
eess.AS2024
SeMaScore : a new evaluation metric for automatic speech recognition tasks
Zitha Sasindran, Harsha Yelchuri, T. V. Prabhakar
In this study, we present SeMaScore, generated using a segment-wise mapping and scoring algorithm that serves as an evaluation metric for automatic speech recognition tasks. SeMaSc…