6 papers
Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models
Nikita Kachaev, Andrey Moskalenko, Matvey Skripkin +10
Embodied Vision-Language-Action (VLA) models are typically obtained by fine-tuning powerful pretrained VLMs on robotics data, yet it is unclear how much commonsense and factual kno…
Listener-Rewarded Thinking in VLMs for Image Preferences
Alexander Gambashidze, Li Pengyi, Matvey Skripkin +5
Training robust and generalizable reward models for human visual preferences is essential for aligning text-to-image and text-to-video generative models with human intent. However,…
Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and Sentences
Dmitrii Korzh, Dmitrii Tarasov, Artyom Iudin +6
Conversion of spoken mathematical expressions is a challenging task that involves transcribing speech into a strictly structured symbolic representation while addressing the ambigu…
Simple Vision-Language Math Reasoning via Rendered Text
Matvey Skripkin, Elizaveta Goncharova, Andrey Kuznetsov
We present a lightweight yet effective pipeline for training vision-language models to solve math problems by rendering LaTeX encoded equations into images and pairing them with st…
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Pengyi Li, Matvey Skripkin, Alexander Zubrey +2
Large language models (LLMs) excel at reasoning, yet post-training remains critical for aligning their behavior with task goals. Existing reinforcement learning (RL) methods often…
MOVE: A Mixture-of-Vision-Encoders Approach for Domain-Focused Vision-Language Processing
Matvey Skripkin, Elizaveta Goncharova, Dmitrii Tarasov +1
Multimodal language models (MLMs) integrate visual and textual information by coupling a vision encoder with a large language model through the specific adapter. While existing app…