papers

Publications (5)

cs.CL2026

LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes

Michael Solodko, Steven Gong, Guangwei Yu +3

LakeQuest is a human‑validated benchmark of 9,846 question‑answer pairs for evaluating end‑to‑end retrieval and synthesis over heterogeneous data lakes across AI/ML metadata, retai…

#question answering#data lakes#retrieval-augmented generation#multimodal
cs.AI2026

CUA-Skill: Develop Skills for Computer Using Agent

Tianyi Chen, Yinheng Li, Michael Solodko +12

Computer-Using Agents (CUAs) aim to autonomously operate computer systems to complete real-world tasks. However, existing agentic systems remain difficult to scale and lag behind h…

cs.AI2026

ScreenSearch: Uncertainty-Aware OS Exploration

Michael Solodko, Justin Wagle

Desktop GUI agents operate under partial observability: visually similar screens can correspond to different underlying workflow states, so locally plausible actions can lead to sh…

cs.CL2025

Self-reflecting Large Language Models: A Hegelian Dialectical Approach

Sara Abdali, Can Goksen, Michael Solodko +4

In this paper, we introduce a self-reflection framework for Large Language Models (LLMs) grounded in the Hegelian Dialectic, a philosophical method in which an initial proposition…

cs.CL2025

AppSelectBench: Application-Level Tool Selection Benchmark

Tianyi Chen, Michael Solodko, Sen Wang +14

Computer Using Agents (CUAs) are increasingly equipped with external tools, enabling them to perform complex and realistic tasks. For CUAs to operate effectively, application selec…