4 papers
Describe-Then-Act: Proactive Agent Steering via Distilled Language-Action World Models
Massimiliano Pappa, Luca Romani, Valentino Sacco +5
Deploying safety-critical agents requires anticipating the consequences of actions before they are executed. While world models offer a paradigm for this proactive foresight, curre…
MonSTeR: a Unified Model for Motion, Scene, Text Retrieval
Luca Collorone, Matteo Gioia, Massimiliano Pappa +5
Intention drives human movement in complex environments, but such movement can only happen if the surrounding context supports it. Despite the intuitive nature of this mechanism, e…
ANTHROPOS-V: benchmarking the novel task of Crowd Volume Estimation
Luca Collorone, Stefano D'Arrigo, Massimiliano Pappa +3
We introduce the novel task of Crowd Volume Estimation (CVE), defined as the process of estimating the collective body volume of crowds using only RGB images. Besides event managem…
MoDiPO: text-to-motion alignment via AI-feedback-driven Direct Preference Optimization
Massimiliano Pappa, Luca Collorone, Giovanni Ficarra +2
Diffusion Models have revolutionized the field of human motion generation by offering exceptional generation quality and fine-grained controllability through natural language condi…