9 papers
The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection
Wojciech Zarzecki, Jan DubiÅski, Sebastian Cygert
Benchmark contamination, where evaluation examples appear in a model's training data, threatens the validity of LLM assessment. Statistical tools for detecting training-data member…
Beyond Classification: Dynamic Adapter Routing for Continual Multimodal Retrieval
Alicja Dobrzeniecka, Filip Szatkowski, Sebastian Cygert +2
While retrieval is a core function of vision-language models, continually updating these models for retrieval tasks remains critically underexplored. Existing work often approaches…
Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics
Maciej ChrabÄ szcz, Aleksander Szymczyk, Marcin Sendera +2
Large Reasoning Models (LRMs) introduce new opportunities for safety monitoring through their Chain of Thought (CoT) reasoning. However, CoT is not always faithful to the model's f…
Efficient Multi-Source Knowledge Transfer by Model Merging
Marcin Osial, Bartosz Wójcik, Bartosz ZieliÅski +1
While transfer learning is an effective strategy, it often overlooks the opportunity to leverage knowledge from numerous available models online. Addressing this multi-source trans…
Subspace Optimization for Backpropagation-Free Continual Test-Time Adaptation
Damian Sójka, Sebastian Cygert, Marc Masana
We introduce PACE, a backpropagation-free continual test-time adaptation system that directly optimizes the affine parameters of normalization layers. Existing derivative-free appr…
Annotation-Efficient Vision-Language Model Adaptation to the Polish Language Using the LLaVA Framework
Grzegorz Statkiewicz, Alicja Dobrzeniecka, Karolina Seweryn +5
Most vision-language models (VLMs) are trained on English-centric data, limiting their performance in other languages and cultural contexts. This restricts their usability for non-…