4 papers · 1 filter
Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation
Kaichen Zhang, Wei Huang, Keming Wu +2
Existing streaming multimodal models process observations incrementally but still follow a turn-based prefill-then-decode pattern, making them non-duplex: new observations cannot n…
TaskGround: Structured Executable Task Inference for Full-Scene Household Reasoning
ZhiYuan Feng, Yu Deng, Ruichuan An +11
In real home deployments, household agents must often operate from a complete household scene and a situated household request, rather than from a clean task specification. Such re…
Dr. Bench: A Multidimensional Evaluation for Deep Research Agents, from Answers to Reports
Yang Yao, Yixu Wang, Yuxuan Zhang +9
As an embodiment of intelligence evolution toward interconnected architectures, Deep Research Agents (DRAs) systematically exhibit the capabilities in task decomposition, cross-sou…
OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe
Kaichen Zhang, Keming Wu, Zuhao Yang +6
Recent advancements in large reasoning models have fueled growing interest in extending such capabilities to multimodal domains. However, despite notable progress in visual reasoni…