From the 1 of 5 linked papers with an AI index.
5 papers
MMLDSum-LLM: Multimodal Long-Document Summarization with Visual-Alignment and Keyword-Aware
Xianpeng Zhang, Jiahua Yang, Dongyu Chen +7
The paper presents a benchmark for multimodal long-document summarization and a two-stage training framework (MMLDSum-LLM) that incorporates visual-alignment and keyword-aware loss…
When Models Judge Themselves: Unsupervised Self-Evolution for Multimodal Reasoning
Zhengxian Wu, Kai Shi, Chuanrui Zhang +10
Recent progress in multimodal large language models has led to strong performance on reasoning tasks, but these improvements largely rely on high-quality annotated data or teacher-…
DaMo: Data Mixing Optimizer in Fine-tuning Multimodal LLMs for Mobile Phone Agents
Kai Shi, Jun Yang, Ni Yang +6
Mobile Phone Agents (MPAs) have emerged as a promising research direction due to their broad applicability across diverse scenarios. While Multimodal Large Language Models (MLLMs)…
ReviewInstruct: A Review-Driven Multi-Turn Conversations Generation Method for Large Language Models
Jiangxu Wu, Cong Wang, TianHuang Su +10
The effectiveness of large language models (LLMs) in conversational AI is hindered by their reliance on single-turn supervised fine-tuning (SFT) data, which limits contextual coher…
Binary Neural Networks for Large Language Model: A Survey
Liangdong Liu, Zhitong Zheng, Cong Wang +2
Large language models (LLMs) have wide applications in the field of natural language processing(NLP), such as GPT-4 and Llama. However, with the exponential growth of model paramet…