5 papers
ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP
Sicheng Zhang, Muzammal Naseer, Binzhu Xie +5
CLIP and its variants are widely adopted visual backbones in multimodal systems, but their pretraining remains dominated by descriptive image-text alignment. As downstream applicat…
SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs
Mohamad Alansari, Naufal Suryanto, Divya Velayudhan +3
Multimodal large language models (MLLMs) have advanced from image-level reasoning to pixel-level grounding, but extending these capabilities to videos remains challenging as models…
RedSage: A Cybersecurity Generalist LLM
Naufal Suryanto, Muzammal Naseer, Pengfei Li +5
Cybersecurity operations demand assistant LLMs that support diverse workflows without exposing sensitive data. Existing solutions either rely on proprietary APIs with privacy risks…
CIPHER: Cybersecurity Intelligent Penetration-testing Helper for Ethical Researcher
Derry Pratama, Naufal Suryanto, Andro Aprila Adiputra +4
Penetration testing, a critical component of cybersecurity, typically requires extensive time and effort to find vulnerabilities. Beginners in this field often benefit from collabo…
Cityscape-Adverse: Benchmarking Robustness of Semantic Segmentation with Realistic Scene Modifications via Diffusion-Based Image Editing
Naufal Suryanto, Andro Aprila Adiputra, Ahmada Yusril Kadiptya +4
Recent advancements in generative AI, particularly diffusion-based image editing, have enabled the transformation of images into highly realistic scenes using only text instruction…