2 papers
cs.AI2026
Skill-based Agentic Evaluation for Real-time Data Science Tasks
Aniruddha Tamhane, Raghavendra Addanki, Ayushi Aggarwal +4
We present a framework for evaluating data-science agents on live, continuously updated data using executable ground truth and format-agnostic factoid scoring. Consider this exampl…
cs.LG2026
Human-LLM Collaborative Feature Engineering for Tabular Data
Zhuoyan Li, Aditya Bansal, Jinzhao Li +8
Large language models (LLMs) are increasingly used to automate feature engineering in tabular learning. Given task-specific information, LLMs can propose diverse feature transforma…