most citedOmniFusion Technical Report

2 citations · 2 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CV2025

Inverting Black-Box Face Recognition Systems via Zero-Order Optimization in Eigenface Space

Anton Razzhigaev, Matvey Mikhalchuk, Klim Kireev +3

Reconstructing facial images from black-box recognition models poses a significant privacy threat. While many methods require access to embeddings, we address the more challenging…

cs.CV2025

Image Reconstruction as a Tool for Feature Analysis

Eduard Allakhverdov, Dmitrii Tarasov, Elizaveta Goncharova +1

Vision encoders are increasingly used in modern applications, from vision-only models to multimodal systems such as vision-language models. Despite their remarkable success, it rem…

cs.CV2025

When Less is Enough: Adaptive Token Reduction for Efficient Image Representation

Eduard Allakhverdov, Elizaveta Goncharova, Andrey Kuznetsov

Vision encoders typically generate a large number of visual tokens, providing information-rich representations but significantly increasing computational demands. This raises the q…

cs.CL2024

Addressing Hallucinations in Language Models with Knowledge Graph Embeddings as an Additional Modality

Viktoriia Chekalina, Anton Razzhigaev, Elizaveta Goncharova +1

In this paper we present an approach to reduce hallucinations in Large Language Models (LLMs) by incorporating Knowledge Graphs (KGs) as an additional modality. Our method involves…

cs.CV20242 cited

OmniFusion Technical Report

Elizaveta Goncharova, Anton Razzhigaev, Matvey Mikhalchuk +6

Last year, multimodal architectures served up a revolution in AI-based approaches and solutions, extending the capabilities of large language models (LLM). We propose an \textit{Om…