activity
20242026
collaborators

6 papers

cs.AI2026

Towards Knowledgeable Deep Research: Framework and Benchmark

Wenxuan Liu, Zixuan Li, Long Bai +13

Deep Research (DR) requires LLM agents to autonomously perform multi-step information seeking, processing, and reasoning to generate comprehensive reports. In contrast to existing…

cs.CV2026

ROMA: Real-time Omni-Multimodal Assistant with Interactive Streaming Understanding

Xueyun Tian, Wei Li, Bingbing Xu +3

Recent Omni-multimodal Large Language Models show promise in unified audio, vision, and text modeling. However, streaming audio-video understanding remains challenging, as existing…

cs.CV2025

MIGE: Mutually Enhanced Multimodal Instruction-Based Image Generation and Editing

Xueyun Tian, Wei Li, Bingbing Xu +3

Despite significant progress in diffusion-based image generation, subject-driven generation and instruction-based editing remain challenging. Existing methods typically treat them…

cs.AI2025

KnowCoder-V2: Deep Knowledge Analysis

Zixuan Li, Wenxuan Liu, Long Bai +13

Deep knowledge analysis tasks always involve the systematic extraction and association of knowledge from large volumes of data, followed by logical reasoning to discover insights.…

cs.CL2024

Fact-Level Confidence Calibration and Self-Correction

Yige Yuan, Bingbing Xu, Hexiang Tan +5

Confidence calibration in LLMs, i.e., aligning their self-assessed confidence with the actual accuracy of their responses, enabling them to self-evaluate the correctness of their o…

cs.CL2024

Unlocking the Power of Large Language Models for Entity Alignment

Xuhui Jiang, Yinghan Shen, Zhichao Shi +6

Entity Alignment (EA) is vital for integrating diverse knowledge graph (KG) data, playing a crucial role in data-driven AI applications. Traditional EA methods primarily rely on co…