6 papers
Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills
Gaytri Jena, Kapil Wanaskar, Vinija Jain +3
Robot learning is splitting into two bets: policies that bake competence into frozen weights (vision-language-action, or VLA, models), and agents that write and refine their own ex…
FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds
Kapil Wanaskar, Gaytri Jena, Aman Chadha +3
World models have attracted significant attention for their ability to capture and predict the structure and dynamics of the physical world. In this emerging landscape, Joint Embed…
A Comprehensive Dataset for Human vs. AI Generated Image Detection
Rajarshi Roy, Ashhar Aziz, Shashwat Bajpai +17
Multimodal generative AI systems like Stable Diffusion, DALL-E, and MidJourney have fundamentally changed how synthetic images are created. These tools drive innovation but also en…
A Comprehensive Dataset for Human vs. AI Generated Text Detection
Rajarshi Roy, Gurpreet Singh, Ashhar Aziz +17
The rapid advancement of large language models (LLMs) has led to increasingly human-like AI-generated text, raising concerns about content authenticity, misinformation, and trustwo…
ECLIPTICA -- A Framework for Switchable LLM Alignment via CITA - Contrastive Instruction-Tuned Alignment
Kapil Wanaskar, Gaytri Jena, Vinija Jain +2
Alignment in large language models (LLMs) is still largely static: after training, the policy is frozen. DPO, GRPO methods typically imprint one behavior into the weights, leaving…
Multimodal Benchmarking and Recommendation of Text-to-Image Generation Models
Kapil Wanaskar, Gaytri Jena, Magdalini Eirinaki
This work presents an open-source unified benchmarking and evaluation framework for text-to-image generation models, with a particular focus on the impact of metadata augmented pro…