works on

From the 1 of 7 linked papers with an AI index.

activity
20242026
collaborators

7 papers

cs.CV2026

LaViDa: A Large Diffusion Language Model for Multimodal Understanding

Shufan Li, Konstantinos Kallidromitis, Hritik Bansal +7

LaViDa introduces a diffusion-based vision-language model that combines a vision encoder with discrete diffusion to enable fast parallel decoding and controllable multimodal genera…

cs.RO2026

Contrastive Action-Image Pre-training for Visuomotor Control

Yuvan Sharma, Dantong Niu, Anirudh Pai +16

Existing vision encoders for robotics face a fundamental bottleneck: robotic datasets lack the scale necessary for large-scale pre-training. Prior work circumvents this data scarci…

cs.AI2025

MobileWorldBench: Towards Semantic World Modeling For Mobile Agents

Shufan Li, Konstantinos Kallidromitis, Akash Gokul +3

World models have shown great utility in improving the task performance of embodied agents. While prior work largely focuses on pixel-space world models, these approaches face prac…

cs.MM2025

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows

Shufan Li, Konstantinos Kallidromitis, Akash Gokul +4

We introduce OmniFlow, a novel generative model designed for any-to-any generation tasks such as text-to-image, text-to-audio, and audio-to-image synthesis. OmniFlow advances the r…

cs.CV2025

Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context Reflection

Shufan Li, Konstantinos Kallidromitis, Akash Gokul +4

The predominant approach to advancing text-to-image generation has been training-time scaling, where larger models are trained on more data using greater computational resources. W…

cs.CV2024

SegLLM: Multi-round Reasoning Segmentation

XuDong Wang, Shaolun Zhang, Shufan Li +5

We present SegLLM, a novel multi-round interactive reasoning segmentation model that enhances LLM-based segmentation by exploiting conversational memory of both visual and textual…