From the 3 of 11 linked papers with an AI index.
11 papers
The GEST-Engine: From Event Graphs to Synthetic Video. A Full Technical Report
Nicolae Cudlenco, Mihai Masala, Marius Leordeanu
The paper presents the GEST-Engine, a system that transforms natural‑language descriptions into fully annotated synthetic video by using an explicit graph‑based world model execute…
GTASA: Ground Truth Annotations for Spatiotemporal Analysis, Evaluation and Training of Video Models
Nicolae Cudlenco, Mihai Masala, Marius Leordeanu
The paper presents GEST-Engine, an open‑source game‑engine system that can render videos with exhaustive, deterministic ground‑truth annotations (entity states, camera parameters,…
Authoring for Living Worlds: Tool-Constrained LLM Agents for Executable Multi-Actor Scenarios
Nicolae Cudlenco, Mihai Masala, Marius Leordeanu
Authoring a multi-actor scenario for a living 3D world, where every action changes its state, and each action's validity depends on the state accumulated before it, demands the fre…
"ÃnÅ£elegi RomâneÅte?'' A Recipe for Romanian Vision-Language Models
Mihai Masala, Marius Leordeanu, Mihai Dascalu +1
Vision-Language Models (VLMs) largely follow the text-only LLM trajectory, excelling on English benchmarks but sharply degrading on low-resource languages, where neither large-scal…
idSCD: Identifying Training Datasets through Semantic Correlation Descriptors
Andrada Gobeaja, Ionut Hodoroaga, Elena Burceanu +1
Can a dataset be recognized from the spurious correlations it induces during training? We argue that datasets leave dataset-specific traces in a model's learned semantic correlatio…
Learning from Random Subspace Exploration: Generalized Test-Time Augmentation with Self-supervised Distillation
Andrei Jelea, Ahmed Nabil Belbachir, Marius Leordeanu
We introduce Generalized Test-Time Augmentation (GTTA), a highly effective method for improving the performance of a trained model, which unlike other existing Test-Time Augmentati…