3 papers
cs.CL2026
Emergent Inference-Time Semantic Contamination via In-Context Priming
Marcin Abram
Recent work has shown that fine-tuning large language models (LLMs) on insecure code or culturally loaded numeric codes can induce emergent misalignment, causing models to produce…
cs.CY2026
Toward Evaluation Frameworks for Multi-Agent Scientific AI Systems
Marcin Abram
We analyze the challenges of benchmarking scientific (multi)-agentic systems, including the difficulty of distinguishing reasoning from retrieval, the risks of data/model contamina…
cs.CL2024
Why does in-context learning fail sometimes? Evaluating in-context learning on open and closed questions
Xiang Li, Haoran Tang, Siyu Chen +3
We measure the performance of in-context learning as a function of task novelty and difficulty for open and closed questions. For that purpose, we created a novel benchmark consist…