most citedUnlocking the Potential of Text-to-Image Diffusion with PAC-Bayesian Theory

2 citations · 3 across the 4 of their papers we have counts for

collaborators

10 papers

cs.AI2025

MobileWorldBench: Towards Semantic World Modeling For Mobile Agents

Shufan Li, Konstantinos Kallidromitis, Akash Gokul +3

World models have shown great utility in improving the task performance of embodied agents. While prior work largely focuses on pixel-space world models, these approaches face prac…

cs.CV2025

Accelerating Inference of Masked Image Generators via Reinforcement Learning

Pranav Subbaraman, Shufan Li, Siyan Zhao +1

Masked Generative Models (MGM)s demonstrate strong capabilities in generating high-fidelity images. However, they need many sampling steps to create high-quality generations, resul…

cs.LG2025

From Masks to Worlds: A Hitchhiker's Guide to World Models

Jinbin Bai, Yu Lei, Hecong Wu +7

This is not a typical survey of world models; it is a guide for those who want to build worlds. We do not aim to catalog every paper that has ever mentioned a ``world model". Inste…

cs.LG2025

PhysiX: A Foundation Model for Physics Simulations

Tung Nguyen, Arsh Koneru, Shufan Li +1

Foundation models have achieved remarkable success across video, image, and language domains. By scaling up the number of parameters and training datasets, these models acquire gen…

cs.CL20251 cited

Mercury: Ultra-Fast Language Models Based on Diffusion

Inception Labs, Samar Khanna, Siddhant Kharbanda +10

We present Mercury, a new generation of commercial-scale large language models (LLMs) based on diffusion. These models are parameterized via the Transformer architecture and traine…

cs.CL2025

PredGen: Accelerated Inference of Large Language Models through Input-Time Speculation for Real-Time Speech Interaction

Shufan Li, Aditya Grover

Large Language Models (LLMs) are widely used in real-time voice chat applications, typically in combination with text-to-speech (TTS) systems to generate audio responses. However,…