5 papers
The Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMs
Ahmed Oumar El-Shangiti, Abzal Nurgazy, Hilal AlQuabeh +2
Despite strong performance on many multimodal tasks, vision-language models (VLMs) still struggle with basic object counting. We investigate whether this reflects missing internal…
The FIL Hypothesis: Inductive Biases Help with Kernel Engineering
Nikolai Rozanov, Subhabrata Dutta, Preslav Nakov +1
The Bitter Lesson, which posits that general-purpose methods that scale with computation and data ultimately outperform those with built-in human knowledge, has become a dominant p…
ThinkBooster: A Unified Framework for Seamless Test-Time Scaling of LLM Reasoning
Vladislav Smirnov, Chieu Nguyen, Sergey Senichev +14
Test-time compute (TTC) scaling has emerged as a powerful paradigm for improving large language model (LLM) reasoning by allocating additional compute during inference, e.g., via m…
Fine-tuning with RAG for Improving LLM Learning of New Skills
Humaid Ibrahim, Nikolai Rozanov, Marek Rei
Large language model (LLM) agents deployed for multi-step tasks frequently fail in predictable ways: attempting actions with unmet preconditions, issuing redundant commands, or mis…
StateAct: Enhancing LLM Base Agents via Self-prompting and State-tracking
Nikolai Rozanov, Marek Rei
Large language models (LLMs) are increasingly used as autonomous agents, tackling tasks from robotics to web navigation. Their performance depends on the underlying base agent. Exi…