11 papers
Visual Grounding in Zero-Shot Vision-Language Control
J. de Curtò, Dayani Plasencia, Diego Sánchez +1
Vision-language models (VLMs) are increasingly used as zero-shot controllers, but successful trajectories do not necessarily show that decisions are grounded in visual input: simul…
Strategic Evaluation of Planning Strategies for LLM Agents in Cyber-Physical Systems
J. de Curtò, I. de Zarzà
Evaluations of LLM planning agents largely ask whether a task succeeds or a declared plan is followed. In strategic cyber-physical systems, a stronger question is whether the plann…
Collective Intelligence with Foundation Models
J. de Curtò, I. de ZarzÃ
As foundation models grow in scale and diversity, coordinating multiple models into cooperative reasoning systems offers a path toward safer, more reliable AI. This chapter present…
Foundation Models for Automatic CAD Generation
J de Curtò, Victoria Guillén, I. de ZarzÃ
Recent advances in Large Language Models (LLMs) and Vision-Language Models (VLMs) enable the automatic generation of parametric 3D designs from natural-language specifications. Thi…
LLM-Mediated Demand Response Coordination in Smart Microgrids
J. de Curtò, I. de ZarzÃ
Effective demand response in smart microgrids requires prosumers to cooperate voluntarily under strategic self-interest, a coordination problem structurally equivalent to a repeate…
Language-Conditioned Visual Grounding with CLIP Multilingual
J. de Curtò, Mauro Liz, I. de ZarzÃ
Multilingual vision-language models exhibit systematic performance gaps across languages, but the mechanism remains ambiguous: cross-language divergence could arise from the visual…