8 papers
CrochetBench: Can Vision-Language Models Move from Describing to Doing in Crochet Domain?
Peiyu Li, Xiaobao Huang, Ting Hua +1
The paper introduces CrochetBench, a benchmark that tests vision-language models on their ability to recognize crochet stitches, ground instructions, and generate executable croche…
ASAP: Agent-System Co-Design for Wall-Clock-Centered Auto HPO Research for ML Experiments
Taicheng Guo, Haomin Zhuang, Kehan Guo +4
Hyperparameter Optimization (HPO) is essential for maximizing machine learning model performance, and its core challenge is sample efficiency: finding strong configurations within…
AutoLLMResearch: Training Research Agents for Automating LLM Experiment Configuration - Learning from Cheap, Optimizing Expensive
Taicheng Guo, Nitesh V. Chawla, Olaf Wiest +1
Effectively configuring scalable large language model (LLM) experiments, spanning architecture design, hyperparameter tuning, and beyond, is crucial for advancing LLM research, as…
On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective
Yue Huang, Chujie Gao, Siyuan Wu +63
Generative Foundation Models (GenFMs) have emerged as transformative tools. However, their widespread adoption raises critical concerns regarding trustworthiness across dimensions.…
MolX: Enhancing Large Language Models for Molecular Understanding With A Multi-Modal Extension
Khiem Le, Zhichun Guo, Kaiwen Dong +8
Large Language Models (LLMs) with their strong task-handling capabilities have shown remarkable advancements across a spectrum of fields, moving beyond natural language understandi…
Transaction Categorization with Relational Deep Learning in QuickBooks
Kaiwen Dong, Padmaja Jonnalagedda, Xiang Gao +5
Automatic transaction categorization is crucial for enhancing the customer experience in QuickBooks by providing accurate accounting and bookkeeping. The distinct challenges in thi…