7 papers
Signature-Guided Capacity Occupancy for Dense Expert Merging
Lingching Tung, Chi-Jui Kim, Beicheng Xu +2
Dense expert merging combines domain-specialized language models into one single checkpoint, typically by admitting task-vector support in weight space. However, this admission is…
MFTune: An Efficient Multi-fidelity Framework for Spark SQL Configuration Tuning
Beicheng Xu, Lingching Tung, Yuchen Wang +2
Apache Spark SQL is a cornerstone of modern big data analytics.However,optimizing Spark SQL performance is challenging due to its vast configuration space and the prohibitive cost…
ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows
Wei Liu, Yang Gu, Xi Yan +5
Table processing-including cleaning, transformation, augmentation, and matching-is a foundational yet error-prone stage in real-world data pipelines. While recent LLM-based approac…
AutoSci: A Memory-Centric Agentic System for the Full Scientific Research Lifecycle
Weitong Qian, Beicheng Xu, Zhongao Xie +16
Scientific research has traditionally been human-intensive, requiring researchers to coordinate literature, ideas, experiments, manuscripts, and review responses across long projec…
CoFEH: LLM-driven Feature Engineering Empowered by Collaborative Bayesian Hyperparameter Optimization
Beicheng Xu, Keyao Ding, Wei Liu +2
Feature Engineering (FE) is pivotal in automated machine learning (AutoML) but remains a bottleneck for traditional methods, which operate within rigid search spaces and lack domai…
Tree-Structured Synergy of Large Language Models and Bayesian Optimization for Efficient CASH
Beicheng Xu, Weitong Qian, Lingching Tung +2
To lower the expertise barrier in machine learning, the AutoML community has focused on the CASH problem, which jointly automates algorithm selection and hyperparameter tuning. Whi…