2 citations · 2 across the 11 of their papers we have counts for
14 papers · 1 filter
Paved with True Intents: Intent-Aware Training Improves LLM Safety Classification Across Training Regimes
Jeremias Ferrao, Niclas Müller-Hof, Iustin Sîrbu +2
We argue that safety classifiers should model user intent as an explicit signal between the prompt and the final label. To study this, we introduce AIMS, a human-annotated dataset…
"ÃnÅ£elegi RomâneÅte?'' A Recipe for Romanian Vision-Language Models
Mihai Masala, Marius Leordeanu, Mihai Dascalu +1
Vision-Language Models (VLMs) largely follow the text-only LLM trajectory, excelling on English benchmarks but sharply degrading on low-resource languages, where neither large-scal…
Semi-Supervised Learning for Large Language Models Safety and Content Moderation
Eduard Stefan Dinuta, Iustin Sirbu, Traian Rebedea
Safety for Large Language Models (LLMs) has been an ongoing research focus since their emergence and is even more relevant nowadays with the increasing capacity of those models. Cu…
Pluralistic Behavior Suite: Stress-Testing Multi-Turn Adherence to Custom Behavioral Policies
Prasoon Varshney, Makesh Narsimhan Sreedhar, Liwei Jiang +2
Large language models (LLMs) are typically aligned to a universal set of safety and usage principles intended for broad public acceptability. Yet, real-world applications of LLMs o…
Improving Romanian LLM Pretraining Data using Diversity and Quality Filtering
Vlad Negoita, Mihai Masala, Traian Rebedea
Large Language Models (LLMs) have recently exploded in popularity, often matching or outperforming human abilities on many tasks. One of the key factors in training LLMs is the ava…
MultiMatch: Multihead Consistency Regularization Matching for Semi-Supervised Text Classification
Iustin Sirbu, Robert-Adrian Popovici, Cornelia Caragea +2
We introduce MultiMatch, a novel semi-supervised learning (SSL) algorithm combining the paradigms of co-training and consistency regularization with pseudo-labeling. At its core, M…