works on

From the 1 of 8 linked papers with an AI index.

activity
20242026
collaborators

8 papers

cs.LG2026

Generalized Fisher-Weighted SVD: Scalable Kronecker-Factored Fisher Approximation for Compressing Large Language Models

Viktoriia Chekalina, Daniil Moskovskiy, Tatiana Matveeva +2

The paper introduces Generalized Fisher-Weighted SVD (GFWSVD), a post‑training compression method for large language models that uses a scalable Kronecker‑factored approximation of…

cs.CL2026

Boosting Self-Consistency with Ranking

Maria Marina, Daniil Moskovskiy, Sergey Pletenev +3

Self-consistency improves large language models by sampling multiple reasoning paths and selecting the most frequent answer, but majority voting often fails to recover correct answ…

cs.CL2026

Leveraging LLM Parametric Knowledge for Fact Checking without Retrieval

Artem Vazhentsev, Maria Marina, Daniil Moskovskiy +8

Trustworthiness is a core research challenge for agentic AI systems built on Large Language Models (LLMs). To enhance trust, natural language claims from diverse sources, including…

cs.CL2025

<think> So let's replace this phrase with insult... </think> Lessons learned from generation of toxic texts with LLMs

Sergey Pletenev, Daniil Moskovskiy, Alexander Panchenko

Modern Large Language Models (LLMs) are excellent at generating synthetic data. However, their performance in sensitive domains such as text detoxification has not received proper…

cs.CL2025

HatePRISM: Policies, Platforms, and Research Integration. Advancing NLP for Hate Speech Proactive Mitigation

Naquee Rizwan, Seid Muhie Yimam, Daryna Dementieva +11

Despite regulations imposed by nations and social media platforms, e.g. (Government of India, 2021; European Parliament and Council of the European Union, 2022), inter alia, hatefu…

cs.CL2025

How Much Knowledge Can You Pack into a LoRA Adapter without Harming LLM?

Sergey Pletenev, Maria Marina, Daniil Moskovskiy +4

The performance of Large Language Models (LLMs) on many tasks is greatly limited by the knowledge learned during pre-training and stored in the model's parameters. Low-rank adaptat…