collaborators

7 papers

cs.RO2025

BLAZER: Bootstrapping LLM-based Manipulation Agents with Zero-Shot Data Generation

Rocktim Jyoti Das, Harsh Singh, Diana Turmakhan +5

Scaling data and models has played a pivotal role in the remarkable progress of computer vision and language. Inspired by these domains, recent efforts in robotics have similarly f…

cs.CV2025

Towards Reliable Identification of Diffusion-based Image Manipulations

Alex Costanzino, Woody Bayliss, Juil Sock +5

Changing facial expressions, gestures, or background details may dramatically alter the meaning conveyed by an image. Notably, recent advances in diffusion models greatly improve t…

cs.LG2025

PSyDUCK: Training-Free Steganography for Latent Diffusion

Aqib Mahfuz, Georgia Channing, Mark van der Wilk +3

Recent advances in generative AI have opened promising avenues for steganography, which can securely protect sensitive information for individuals operating in hostile environments…

cs.CV2024

AlignGuard: Scalable Safety Alignment for Text-to-Image Generation

Runtao Liu, I Chieh Chen, Jindong Gu +6

Text-to-image (T2I) models are widespread, but their limited safety guardrails expose end users to harmful content and potentially allow for model misuse. Current safety measures a…

cs.CV2024

Video Motion Transfer with Diffusion Transformers

Alexander Pondaven, Aliaksandr Siarohin, Sergey Tulyakov +2

We propose DiTFlow, a method for transferring the motion of a reference video to a newly synthesized one, designed specifically for Diffusion Transformers (DiT). We first process t…

cs.LG2024

MALT: Improving Reasoning with Multi-Agent LLM Training

Sumeet Ramesh Motwani, Chandler Smith, Rocktim Jyoti Das +6

Large Language Models (LLMs) often produce answers with a single chain-of-thought, which restricts their ability to explore reasoning paths or self-correct flawed outputs in comple…