3 papers
cs.CL2026
Eta Given Delta: Defining LLM Tool Efficiency With Marginal Tool Utility
Nyx Iskandar
This paper introduces tool efficiency, a new quantitative metric to evaluate the rate of useful tool calls in an LLM agent trajectory. To ensure that tool efficiency is well-define…
cs.CL2025
A Matter of Representation: Towards Graph-Based Abstract Code Generation
Nyx Iskandar, Hisham Bedri, Andy Tsen
Most large language models (LLMs) today excel at generating raw, sequential code with minimal abstractions and custom structures. However, there has been little work on graph-based…
cs.SE2025
Measuring What Matters: A Framework for Evaluating Safety Risks in Real-World LLM Applications
Jia Yi Goh, Shaun Khoo, Nyx Iskandar +3
Most safety testing efforts for large language models (LLMs) today focus on evaluating foundation models. However, there is a growing need to evaluate safety at the application lev…