collaborators

5 papers

cs.CV2026

Reasoning in Computer Vision: Taxonomy, Models, Tasks, and Methodologies

Ayushman Sarkar, Zhenyu Yu, Mohd Yamani Idna Idris

Visual reasoning matters for many computer vision tasks that go beyond surface-level object detection and classification. Despite progress in relational, symbolic, temporal, causal…

cs.CV2026

DeCorStory: Gram-Schmidt Prompt Embedding Decorrelation for Consistent Storytelling

Ayushman Sarkar, Zhenyu Yu, Mohd Yamani Idna Idris

Maintaining visual and semantic consistency across frames is a key challenge in text-to-image storytelling. Existing training-free methods, such as One-Prompt-One-Story, concatenat…

cs.CV2026

StoryState: Agent-Based State Control for Consistent and Editable Storybooks

Ayushman Sarkar, Zhenyu Yu, Wei Tang +3

Large multimodal models have enabled one-click storybook generation, where users provide a short description and receive a multi-page illustrated story. However, the underlying sto…

cs.CV2026

ReDiStory: Region-Disentangled Diffusion for Consistent Visual Story Generation

Ayushman Sarkar, Zhenyu Yu, Chu Chen +3

Generating coherent visual stories requires maintaining subject identity across multiple images while preserving frame-specific semantics. Recent training-free methods concatenate…

cs.CV2025

STAC: Leveraging Spatio-Temporal Data Associations For Efficient Cross-Camera Streaming and Analytics

Ragini Gupta, Lingzhi Zhao, Jiaxi Li +5

In IoT based distributed network of cameras, real-time multi-camera video analytics is challenged by high bandwidth demands and redundant visual data, creating a fundamental tensio…