2 citations · 2 across the 3 of their papers we have counts for
4 papers
ROMA: Real-time Omni-Multimodal Assistant with Interactive Streaming Understanding
Xueyun Tian, Wei Li, Bingbing Xu +3
Recent Omni-multimodal Large Language Models show promise in unified audio, vision, and text modeling. However, streaming audio-video understanding remains challenging, as existing…
KnowCoder-V2: Deep Knowledge Analysis
Zixuan Li, Wenxuan Liu, Long Bai +13
Deep knowledge analysis tasks always involve the systematic extraction and association of knowledge from large volumes of data, followed by logical reasoning to discover insights.…
MIGE: Mutually Enhanced Multimodal Instruction-Based Image Generation and Editing
Xueyun Tian, Wei Li, Bingbing Xu +3
Despite significant progress in diffusion-based image generation, subject-driven generation and instruction-based editing remain challenging. Existing methods typically treat them…
Fact-Level Confidence Calibration and Self-Correction
Yige Yuan, Bingbing Xu, Hexiang Tan +5
Confidence calibration in LLMs, i.e., aligning their self-assessed confidence with the actual accuracy of their responses, enabling them to self-evaluate the correctness of their o…