From the 1 of 3 linked papers with an AI index.
3 papers
cs.AI2026
Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks
Jeff Mohl, Nelson Gardner-Challis, Magda Dubois +6
The paper presents automated AI scanners that analyze benchmark transcripts to detect validity flaws such as ground‑truth leakage, tool failures, guessing vulnerabilities, and ambi…
cs.LG2026
Subspace Optimization for Backpropagation-Free Continual Test-Time Adaptation
Damian Sójka, Sebastian Cygert, Marc Masana
We introduce PACE, a backpropagation-free continual test-time adaptation system that directly optimizes the affine parameters of normalization layers. Existing derivative-free appr…
cs.LG2024
Realistic Evaluation of Test-Time Adaptation Algorithms: Unsupervised Hyperparameter Selection
Sebastian Cygert, Damian Sójka, Tomasz TrzciÅski +1
Test-Time Adaptation (TTA) has recently emerged as a promising strategy for tackling the problem of machine learning model robustness under distribution shifts by adapting the mode…