2 papers
cs.CV2025
Reconstruction-Driven Multimodal Representation Learning for Automated Media Understanding
Yassir Benhammou, Suman Kalyan, Sujay Kumar
Broadcast and media organizations increasingly rely on artificial intelligence to automate the labor-intensive processes of content indexing, tagging, and metadata generation. Howe…
cs.CV2025
Zero-Shot, But at What Cost? Unveiling the Hidden Overhead of MILS's LLM-CLIP Framework for Image Captioning
Yassir Benhammou, Alessandro Tiberio, Gabriel Trautmann +1
MILS (Multimodal Iterative LLM Solver) is a recently published framework that claims "LLMs can see and hear without any training" by leveraging an iterative, LLM-CLIP based approac…