From the 3 of 9 linked papers with an AI index.
9 papers
The GEST-Engine: From Event Graphs to Synthetic Video. A Full Technical Report
Nicolae Cudlenco, Mihai Masala, Marius Leordeanu
The paper presents the GEST-Engine, a system that transforms natural‑language descriptions into fully annotated synthetic video by using an explicit graph‑based world model execute…
GTASA: Ground Truth Annotations for Spatiotemporal Analysis, Evaluation and Training of Video Models
Nicolae Cudlenco, Mihai Masala, Marius Leordeanu
The paper presents GEST-Engine, an open‑source game‑engine system that can render videos with exhaustive, deterministic ground‑truth annotations (entity states, camera parameters,…
Authoring for Living Worlds: Tool-Constrained LLM Agents for Executable Multi-Actor Scenarios
Nicolae Cudlenco, Mihai Masala, Marius Leordeanu
Authoring a multi-actor scenario for a living 3D world, where every action changes its state, and each action's validity depends on the state accumulated before it, demands the fre…
"ÃnÅ£elegi RomâneÅte?'' A Recipe for Romanian Vision-Language Models
Mihai Masala, Marius Leordeanu, Mihai Dascalu +1
Vision-Language Models (VLMs) largely follow the text-only LLM trajectory, excelling on English benchmarks but sharply degrading on low-resource languages, where neither large-scal…
Improving Romanian LLM Pretraining Data using Diversity and Quality Filtering
Vlad Negoita, Mihai Masala, Traian Rebedea
Large Language Models (LLMs) have recently exploded in popularity, often matching or outperforming human abilities on many tasks. One of the key factors in training LLMs is the ava…
From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach
Mihai Masala, Marius Leordeanu
The task of describing video content in natural language is commonly referred to as video captioning. Unlike conventional video captions, which are typically brief and widely avail…