7 papers
Pulling The REINS: Training-Free Safety Alignment of Video Diffusion Models via Representation Steering
Rohit Kundu, Arindam Dutta, Sarosij Bose +2
Open-weight video diffusion models can generate photorealistic unsafe content, from violence to misinformation, yet existing defenses either require expensive safety fine-tuning th…
Visual Alignment of Medical Vision-Language Models for Grounded Radiology Report Generation
Sarosij Bose, Ravi K. Rajendran, Biplob Debnath +3
Radiology Report Generation (RRG) is a critical step toward automating healthcare workflows, facilitating accurate patient assessments, and reducing the workload of medical profess…
Uncertainty-Aware Diffusion Guided Refinement of 3D Scenes
Sarosij Bose, Arindam Dutta, Sayak Nag +4
Reconstructing 3D scenes from a single image is a fundamentally ill-posed task due to the severely under-constrained nature of the problem. Consequently, when the scene is rendered…
VOccl3D: A Video Benchmark Dataset for 3D Human Pose and Shape Estimation under real Occlusions
Yash Garg, Saketh Bachu, Arindam Dutta +5
Human pose and shape (HPS) estimation methods have been extensively studied, with many demonstrating high zero-shot performance on in-the-wild images and videos. However, these met…
Conformal Prediction and MLLM aided Uncertainty Quantification in Scene Graph Generation
Sayak Nag, Udita Ghosh, Calvin-Khang Ta +3
Scene Graph Generation (SGG) aims to represent visual scenes by identifying objects and their pairwise relationships, providing a structured understanding of image content. However…
Leveraging Synthetic Adult Datasets for Unsupervised Infant Pose Estimation
Sarosij Bose, Hannah Dela Cruz, Arindam Dutta +3
Human pose estimation is a critical tool across a variety of healthcare applications. Despite significant progress in pose estimation algorithms targeting adults, such developments…