From the 1 of 9 linked papers with an AI index.
9 papers
Explainable and Resource-Efficient Spatial Reasoning in Multimodal LLMs for Decision-Critical Applications
Piyush Jain, Kousik Dasgupta, Rajarshi Roy +1
The paper introduces ByDeWay-V2, a training‑free prompting framework that adds explicit pairwise spatial predicates derived from depth estimation and open‑vocabulary object detecti…
DyaPlex: Full-Duplex Speech-Motion Model for Dyadic Interaction
Koki Nagano, Hongyu Liu, Seonwook Park +9
We present DyaPlex, a streaming, full-duplex speech-and-motion model designed for dyadic interaction. To capture the continuous and reciprocal nature of human communication, this f…
VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents
Amrita Mazumdar, Seonwook Park, Rajarshi Roy +6
Natural human conversation is full-duplex and audio-visual: people simultaneously speak and listen while continuously interpreting and producing nonverbal cues, such as nods, smile…
A Comprehensive Dataset for Human vs. AI Generated Image Detection
Rajarshi Roy, Ashhar Aziz, Shashwat Bajpai +17
Multimodal generative AI systems like Stable Diffusion, DALL-E, and MidJourney have fundamentally changed how synthetic images are created. These tools drive innovation but also en…
A Comprehensive Dataset for Human vs. AI Generated Text Detection
Rajarshi Roy, Gurpreet Singh, Ashhar Aziz +17
The rapid advancement of large language models (LLMs) has led to increasingly human-like AI-generated text, raising concerns about content authenticity, misinformation, and trustwo…
ByDeWay: Boost Your multimodal LLM with DEpth prompting in a Training-Free Way
Rajarshi Roy, Devleena Das, Ankesh Banerjee +3
We introduce ByDeWay, a training-free framework designed to enhance the performance of Multimodal Large Language Models (MLLMs). ByDeWay uses a novel prompting strategy called Laye…