4 papers
Agent psychometrics: Task-level performance prediction in agentic coding benchmarks
Chris Ge, Daria Kryvosheieva, Daniel Fried +2
As the focus in LLM-based coding shifts from static single-step code generation to multi-step agentic interaction with tools and environments, understanding which tasks will challe…
Different types of syntactic agreement recruit the same units within large language models
Daria Kryvosheieva, Andrea de Varda, Evelina Fedorenko +1
Large language models (LLMs) can reliably distinguish grammatical from ungrammatical sentences, but how grammatical knowledge is represented within the models remains an open quest…
Efficient Code Embeddings from Code Generation Models
Daria Kryvosheieva, Saba Sturua, Michael Günther +1
jina-code-embeddings is a novel code embedding model suite designed to retrieve code from natural language queries, perform technical question-answering, and identify semantically…
Controlled Evaluation of Syntactic Knowledge in Multilingual Language Models
Daria Kryvosheieva, Roger Levy
Language models (LMs) are capable of acquiring elements of human-like syntactic knowledge. Targeted syntactic evaluation tests have been employed to measure how well they form gene…