Publications (5)
LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes
Michael Solodko, Steven Gong, Guangwei Yu +3
LakeQuest is a human‑validated benchmark of 9,846 question‑answer pairs for evaluating end‑to‑end retrieval and synthesis over heterogeneous data lakes across AI/ML metadata, retai…
CUA-Skill: Develop Skills for Computer Using Agent
Tianyi Chen, Yinheng Li, Michael Solodko +12
Computer-Using Agents (CUAs) aim to autonomously operate computer systems to complete real-world tasks. However, existing agentic systems remain difficult to scale and lag behind h…
ScreenSearch: Uncertainty-Aware OS Exploration
Michael Solodko, Justin Wagle
Desktop GUI agents operate under partial observability: visually similar screens can correspond to different underlying workflow states, so locally plausible actions can lead to sh…
Self-reflecting Large Language Models: A Hegelian Dialectical Approach
Sara Abdali, Can Goksen, Michael Solodko +4
In this paper, we introduce a self-reflection framework for Large Language Models (LLMs) grounded in the Hegelian Dialectic, a philosophical method in which an initial proposition…
AppSelectBench: Application-Level Tool Selection Benchmark
Tianyi Chen, Michael Solodko, Sen Wang +14
Computer Using Agents (CUAs) are increasingly equipped with external tools, enabling them to perform complex and realistic tasks. For CUAs to operate effectively, application selec…