3 papers
cs.AI2026
Outcome Monitors: Recovery Affordances for Silent Tool Failures
Sugam Panthi, Rabab Abdelfattah
When a tool call times out, the agent sees the failure and can route around it. A cached error page or negative price can instead arrive in the expected format and be consumed as f…
cs.IR2026
Same Ranking, Different Winner: How Scoring Targets Shape LLM Memory Benchmarks
Sugam Panthi, Rabab Abdelfattah
Conversational-memory systems increasingly transform dialogue history into facts, summaries, timelines, and other source-linked descendants, so a single source turn can coexist wit…
cs.AI2026
LLM-as-Judge for Semantic Judging of Powerline Segmentation in UAV Inspection
Akram Hossain, Rabab Abdelfattah, Xiaofeng Wang +1
The deployment of lightweight segmentation models on drones for autonomous power line inspection presents a critical challenge: maintaining reliable performance under real-world co…