6 papers
FlowInOne:Unifying Multimodal Generation as Image-in, Image-out Flow Matching
Junchao Yi, Rui Zhao, Jiahao Tang +7
Multimodal generation has long been dominated by text-driven pipelines where language dictates vision but cannot reason or create within it. We challenge this paradigm by asking wh…
Cross-modal Prompting for Balanced Incomplete Multi-modal Emotion Recognition
Wen-Jue He, Xiaofeng Zhu, Zheng Zhang
Incomplete multi-modal emotion recognition (IMER) aims at understanding human intentions and sentiments by comprehensively exploring the partially observed multi-source data. Altho…
Agentic Meta-Orchestrator for Multi-task Copilots
Xiaofeng Zhu, Yunshen Zhou
Microsoft Copilot suites serve as the universal entry point for various agents skilled in handling important tasks, ranging from assisting a customer with product purchases to dete…
Can LLMs Interpret and Leverage Structured Linguistic Representations? A Case Study with AMRs
Ankush Raut, Xiaofeng Zhu, Maria Leonor Pacheco
This paper evaluates the ability of Large Language Models (LLMs) to leverage contextual information in the form of structured linguistic representations. Specifically, we examine t…
Navigating Semantic Drift in Task-Agnostic Class-Incremental Learning
Fangwen Wu, Lechao Cheng, Shengeng Tang +4
Class-incremental learning (CIL) seeks to enable a model to sequentially learn new classes while retaining knowledge of previously learned ones. Balancing flexibility and stability…
Trustful LLMs: Customizing and Grounding Text Generation with Knowledge Bases and Dual Decoders
Xiaofeng Zhu, Jaya Krishna Mandivarapu
Although people are impressed by the content generation skills of large language models, the use of LLMs, such as ChatGPT, is limited by the domain grounding of the content. The co…