2 citations · 2 across the 8 of their papers we have counts for
5 papers · 1 filter
Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and Sentences
Dmitrii Korzh, Dmitrii Tarasov, Artyom Iudin +6
Conversion of spoken mathematical expressions is a challenging task that involves transcribing speech into a strictly structured symbolic representation while addressing the ambigu…
Listener-Rewarded Thinking in VLMs for Image Preferences
Alexander Gambashidze, Li Pengyi, Matvey Skripkin +5
Training robust and generalizable reward models for human visual preferences is essential for aligning text-to-image and text-to-video generative models with human intent. However,…
MOVE: A Mixture-of-Vision-Encoders Approach for Domain-Focused Vision-Language Processing
Matvey Skripkin, Elizaveta Goncharova, Dmitrii Tarasov +1
Multimodal language models (MLMs) integrate visual and textual information by coupling a vision encoder with a large language model through the specific adapter. While existing app…
UniDet3D: Multi-dataset Indoor 3D Object Detection
Maksim Kolodiazhnyi, Anna Vorontsova, Matvey Skripkin +2
Growing customer demand for smart solutions in robotics and augmented reality has attracted considerable attention to 3D object detection from point clouds. Yet, existing indoor da…
OmniFusion Technical Report
Elizaveta Goncharova, Anton Razzhigaev, Matvey Mikhalchuk +6
Last year, multimodal architectures served up a revolution in AI-based approaches and solutions, extending the capabilities of large language models (LLM). We propose an \textit{Om…