104 papers
GuardPaint:SpeculativeSafetyDecodingforText-to-ImageGeneration
Shreyash Dhoot, Paras Dhiman, Arsh Abbas Naqvi +4
Text-to-image (T2I) diffusion models offer powerful visual generation, but their controllability creates a critical safety challenge: adversarial prompts can steer the denoising tr…
Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills
Gaytri Jena, Kapil Wanaskar, Vinija Jain +3
Robot learning is splitting into two bets: policies that bake competence into frozen weights (vision-language-action, or VLA, models), and agents that write and refine their own ex…
FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds
Kapil Wanaskar, Gaytri Jena, Aman Chadha +3
World models have attracted significant attention for their ability to capture and predict the structure and dynamics of the physical world. In this emerging landscape, Joint Embed…
Khondo: A Multimodal Benchmark for Document Packet Splitting of Bangla Forms
Abu Tyeb Azad, Fahim Ahmed, Ishita Sur Apan +7
Document packets, multiple documents concatenated into a single file, are common in government and administrative workflows, yet splitting them into their constituent documents is…
BaFCo: A Document Understanding Benchmark for Complex Bangla Form Comprehension
Abu Tyeb Azad, Ishita Sur Apan, Fahim Ahmed +8
Document comprehension is a challenging yet impactful task for Multimodal Large Language Models, especially as these systems see growing adoption in real-world, human-centric appli…
Pessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adaptation in Reasoning Models
Subramanyam Sahoo, Aman Chadha, Vinija Jain +1
Conservative offline training is widely advocated as a safe foundation for subsequent online adaptation: if a policy stays close to well-supported behaviour, the argument goes, it…