activity
20242026
most citedOmniFusion Technical Report

2 citations · 2 across the 8 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2025

Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and Sentences

Dmitrii Korzh, Dmitrii Tarasov, Artyom Iudin +6

Conversion of spoken mathematical expressions is a challenging task that involves transcribing speech into a strictly structured symbolic representation while addressing the ambigu…

cs.CV2025

Listener-Rewarded Thinking in VLMs for Image Preferences

Alexander Gambashidze, Li Pengyi, Matvey Skripkin +5

Training robust and generalizable reward models for human visual preferences is essential for aligning text-to-image and text-to-video generative models with human intent. However,…

cs.CV2025

MOVE: A Mixture-of-Vision-Encoders Approach for Domain-Focused Vision-Language Processing

Matvey Skripkin, Elizaveta Goncharova, Dmitrii Tarasov +1

Multimodal language models (MLMs) integrate visual and textual information by coupling a vision encoder with a large language model through the specific adapter. While existing app…

cs.CV2024

UniDet3D: Multi-dataset Indoor 3D Object Detection

Maksim Kolodiazhnyi, Anna Vorontsova, Matvey Skripkin +2

Growing customer demand for smart solutions in robotics and augmented reality has attracted considerable attention to 3D object detection from point clouds. Yet, existing indoor da…

cs.CV20242 cited

OmniFusion Technical Report

Elizaveta Goncharova, Anton Razzhigaev, Matvey Mikhalchuk +6

Last year, multimodal architectures served up a revolution in AI-based approaches and solutions, extending the capabilities of large language models (LLM). We propose an \textit{Om…