10 papers
ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
Hangjie Yuan, Yichen Qian, Zhiwei Tang +21
Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric chall…
Agent-as-a-Router: Agentic Model Routing for Coding Tasks
Pengfei Zhou, Zhiwei Tang, Yixing Ma +8
Real-world users typically have access to multiple Large Language Models (LLMs) from different providers, and these LLMs often excel at distinct domains, yet none dominate all. Con…
BIFE: Better Interaction, Fewer Errors for Minute-Long Video Generation
Zeyu Zhang, Jinyuan Mao, Shuning Chang +5
Long video generation is a critical step toward building realistic world models, requiring both high visual fidelity and long-range interaction consistency. Recent autoregressive d…
Towards 3D-Aware Video Diffusion Models: Render-Free Human Motion Control with Mesh Tokenization
Jingyun Liang, Min Wei, Shikai Li +5
Diffusion models have shown remarkable success in video generation. However, whether such models are truly aware of the 3D structure underlying visual observations, rather than sim…
Inferix: A Block-Diffusion based Next-Generation Inference Engine for World Simulation
Inferix Team, Tianyu Feng, Yizeng Han +13
World models serve as core simulators for fields such as agentic AI, embodied AI, and gaming, capable of generating long, physically realistic, and interactive high-quality videos.…
DyDiT++: Diffusion Transformers with Timestep and Spatial Dynamics for Efficient Visual Generation
Wangbo Zhao, Yizeng Han, Jiasheng Tang +6
Diffusion Transformer (DiT), an emerging diffusion model for visual generation, has demonstrated superior performance but suffers from substantial computational costs. Our investig…