works on

From the 1 of 9 linked papers with an AI index.

collaborators

9 papers

cs.CV2026

Explainable and Resource-Efficient Spatial Reasoning in Multimodal LLMs for Decision-Critical Applications

Piyush Jain, Kousik Dasgupta, Rajarshi Roy +1

The paper introduces ByDeWay-V2, a training‑free prompting framework that adds explicit pairwise spatial predicates derived from depth estimation and open‑vocabulary object detecti…

cs.CV2026

DyaPlex: Full-Duplex Speech-Motion Model for Dyadic Interaction

Koki Nagano, Hongyu Liu, Seonwook Park +9

We present DyaPlex, a streaming, full-duplex speech-and-motion model designed for dyadic interaction. To capture the continuous and reciprocal nature of human communication, this f…

cs.CV2026

VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents

Amrita Mazumdar, Seonwook Park, Rajarshi Roy +6

Natural human conversation is full-duplex and audio-visual: people simultaneously speak and listen while continuously interpreting and producing nonverbal cues, such as nods, smile…

cs.CV2026

A Comprehensive Dataset for Human vs. AI Generated Image Detection

Rajarshi Roy, Ashhar Aziz, Shashwat Bajpai +17

Multimodal generative AI systems like Stable Diffusion, DALL-E, and MidJourney have fundamentally changed how synthetic images are created. These tools drive innovation but also en…

cs.CL2026

A Comprehensive Dataset for Human vs. AI Generated Text Detection

Rajarshi Roy, Gurpreet Singh, Ashhar Aziz +17

The rapid advancement of large language models (LLMs) has led to increasingly human-like AI-generated text, raising concerns about content authenticity, misinformation, and trustwo…

cs.CV2025

ByDeWay: Boost Your multimodal LLM with DEpth prompting in a Training-Free Way

Rajarshi Roy, Devleena Das, Ankesh Banerjee +3

We introduce ByDeWay, a training-free framework designed to enhance the performance of Multimodal Large Language Models (MLLMs). ByDeWay uses a novel prompting strategy called Laye…