4 papers
Large-scale Testing Global Optimization Methods with Black-box Adversarial Attacks
Wojciech Zarzecki, Jarosław Arabas
Existing global optimization benchmark suites are of a moderate size and are based on a small number of analytical functions that date back even to the 1970s. This causes a risk of…
The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection
Wojciech Zarzecki, Jan Dubiński, Sebastian Cygert
Benchmark contamination, where evaluation examples appear in a model's training data, threatens the validity of LLM assessment. Statistical tools for detecting training-data member…
seqme: a Python library for evaluating biological sequence design
Rasmus Møller-Larsen, Adam Izdebski, Jan Olszewski +4
Recent advances in computational methods for designing biological sequences have sparked the development of metrics to evaluate these methods performance in terms of the fidelity o…
FoldSAE: Learning to Steer Protein Folding Through Sparse Representations
Wojciech Zarzecki, Paulina Szymczak, Ewa Szczurek +1
RFdiffusion is a popular and well-established model for generation of protein structures. However, this generative process offers limited insight into its internal representations…