2 papers
cs.CL2026
Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs
Guijin Son, Seungone Kim, Catherine Arnett +73
Following the recent achievement of gold-medal performance on the IMO by frontier LLMs, the community is searching for the next meaningful and challenging target for measuring LLM…
cs.LG2024
CAST: Cluster-Aware Self-Training for Tabular Data via Reliable Confidence
Minwook Kim, Juseong Kim, Ki Beom Kim +1
Tabular data is one of the most widely used data modalities, encompassing numerous datasets with substantial amounts of unlabeled data. Despite this prevalence, there is a notable…