papers

Publications (12)

cs.CL2025

Enhancing Multimodal Continual Instruction Tuning with BranchLoRA

Duzhen Zhang, Yong Ren, Zhong-Zhi Li +5

Multimodal Continual Instruction Tuning (MCIT) aims to finetune Multimodal Large Language Models (MLLMs) to continually align with human intent across sequential tasks. Existing ap…

cs.CL2024

MM-LLMs: Recent Advances in MultiModal Large Language Models

Duzhen Zhang, Yahan Yu, Jiahua Dong +4

In the past year, MultiModal Large Language Models (MM-LLMs) have undergone substantial advancements, augmenting off-the-shelf LLMs to support MM inputs or outputs via cost-effecti…

cs.CV2026

Progressive Multimodal Alignment for Continual Instruction Tuning

Duzhen Zhang, Yahan Yu, Qiaoyi Su +2

The paper proposes Progressive Multimodal Alignment (PMA), a framework that adds expandable expert projectors and a routing mechanism to continually adapt visual-language alignment…

#continual learning#multimodal alignment#large language models#visual-language projection
cs.CL2023

Continual Named Entity Recognition without Catastrophic Forgetting

Duzhen Zhang, Wei Cong, Jiahua Dong +4

Continual Named Entity Recognition (CNER) is a burgeoning area, which involves updating an existing model by incorporating new entity types sequentially. Nevertheless, continual le…

cs.CV2025

Guiding Perception-Reasoning Closer to Human in Blind Image Quality Assessment

Yuan Li, Yahan Yu, Youyuan Lin +3

Humans assess image quality through a perception-reasoning cascade, integrating sensory cues with implicit reasoning to form self-consistent judgments. In this work, we investigate…

cs.CL2025

MedKGent: A Large Language Model Agent Framework for Constructing Temporally Evolving Medical Knowledge Graph

Duzhen Zhang, Zixiao Wang, Zhong-Zhi Li +10

The rapid expansion of medical literature challenges the scalable structuring of domain knowledge. Knowledge Graphs (KGs) offer a solution, yet current construction methods lack ge…

cs.CV2026

Learning Deliberately, Acting Intuitively: Unlocking Test-Time Reasoning in Multimodal LLMs

Yahan Yu, Yuyang Dong, Masafumi Oyamada

Reasoning is essential for large language models (LLMs), especially in complex tasks such as mathematical problem solving. However, multimodal reasoning still faces challenges in m…

cs.CL2025

SpeechIQ: Speech-Agentic Intelligence Quotient Across Cognitive Levels in Voice Understanding by Large Language Models

Zhen Wan, Chao-Han Huck Yang, Yahan Yu +8

We introduce Speech-based Intelligence Quotient (SIQ) as a new form of human cognition-inspired evaluation pipeline for voice understanding large language models, LLM Voice, design…

cs.CV2023

Stroke Extraction of Chinese Character Based on Deep Structure Deformable Image Registration

Meng Li, Yahan Yu, Yi Yang +2

Stroke extraction of Chinese characters plays an important role in the field of character recognition and generation. The most existing character stroke extraction methods focus on…

cs.CL2026

Revisiting Anthropomorphic Reflection Markers in Large Language Model Reasoning

Yahan Yu, Noa Nakanishi, Fei Cheng

Large Language Models (LLMs) often produce explicit reflective traces during complex reasoning, accompanied by anthropomorphic markers such as wait, hmm, and alternatively. Althoug…

cs.CL2025

When Large Language Models Meet Speech: A Survey on Integration Approaches

Zhengdong Yang, Shuichiro Shimizu, Yahan Yu +1

Recent advancements in large language models (LLMs) have spurred interest in expanding their application beyond text-based tasks. A large number of studies have explored integratin…

cs.CL2025

Federated Incremental Named Entity Recognition

Duzhen Zhang, Yahan Yu, Chenxing Li +2

Federated Named Entity Recognition (FNER) boosts model training within each local client by aggregating the model updates of decentralized local clients, without sharing their priv…