3 papers
cs.AI2026
Improving LLM Performance Through Black-Box Online Tuning: A Case for Adding System Specs to Factsheets for Trusted AI
Yonas Atinafu, Henry Lin, Robin Cohen
In this paper, we present a novel black-box online controller that uses only end-to-end measurements over short segments, without internal instrumentation, and hill climbing to max…
cs.AI2026
RewardHackingAgents: Benchmarking Evaluation Integrity for LLM ML-Engineering Agents
Yonas Atinafu, Robin Cohen
LLM agents increasingly perform end-to-end ML engineering tasks where success is judged by a single scalar test metric. This creates a structural vulnerability: an agent can increa…
cs.PF2025
PixLift: Accelerating Web Browsing via AI Upscaling
Yonas Atinafu, Sarthak Malla, HyunSeok Daniel Jang +3
Accessing the internet in regions with expensive data plans and limited connectivity poses significant challenges, restricting information access and economic growth. Images, as a…