activity
20232026
most citedEAIRA: Establishing a Methodology for Evaluating AI Models as Scientific Research Assistants

1 citations · 1 across the 13 of their papers we have counts for

collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2026

FAMOSE: A ReAct Approach to Automated Feature Discovery

Keith Burghardt, Jienan Liu, Sadman Sakib +2

Feature engineering remains a critical yet challenging bottleneck in machine learning, particularly for tabular data, as identifying optimal features from an exponentially large fe…

cs.LG2025

Memory-Efficient LLM Training by Various-Grained Low-Rank Projection of Gradients

Yezhen Wang, Zhouhao Yang, Brian K Chen +4

Building upon the success of low-rank adapter (LoRA), low-rank gradient projection (LoRP) has emerged as a promising solution for memory-efficient fine-tuning. However, existing Lo…

cs.LG2025

Anyprefer: An Agentic Framework for Preference Data Synthesis

Yiyang Zhou, Zhaoyang Wang, Tianle Wang +13

High-quality preference data is essential for aligning foundation models with human values through preference learning. However, manual annotation of such data is often time-consum…

cs.LG2025

ControllableGPT: A Ground-Up Designed Controllable GPT for Molecule Optimization

Xuefeng Liu, Songhao Jiang, Bo Li +1

Large Language Models (LLMs) employ three popular training approaches: Masked Language Models (MLM), Causal Language Models (CLM), and Sequence-to-Sequence Models (seq2seq). Howeve…

cs.LG2023

WordScape: a Pipeline to extract multilingual, visually rich Documents with Layout Annotations from Web Crawl Data

Maurice Weber, Carlo Siebenschuh, Rory Butler +8

We introduce WordScape, a novel pipeline for the creation of cross-disciplinary, multilingual corpora comprising millions of pages with annotations for document layout detection. R…