8 papers
Progressive Multimodal Alignment for Continual Instruction Tuning
Duzhen Zhang, Yahan Yu, Qiaoyi Su +2
The paper proposes Progressive Multimodal Alignment (PMA), a framework that adds expandable expert projectors and a routing mechanism to continually adapt visual-language alignment…
Revisiting Anthropomorphic Reflection Markers in Large Language Model Reasoning
Yahan Yu, Noa Nakanishi, Fei Cheng
Large Language Models (LLMs) often produce explicit reflective traces during complex reasoning, accompanied by anthropomorphic markers such as wait, hmm, and alternatively. Althoug…
Guiding Perception-Reasoning Closer to Human in Blind Image Quality Assessment
Yuan Li, Yahan Yu, Youyuan Lin +3
Humans assess image quality through a perception-reasoning cascade, integrating sensory cues with implicit reasoning to form self-consistent judgments. In this work, we investigate…
SpeechIQ: Speech-Agentic Intelligence Quotient Across Cognitive Levels in Voice Understanding by Large Language Models
Zhen Wan, Chao-Han Huck Yang, Yahan Yu +8
We introduce Speech-based Intelligence Quotient (SIQ) as a new form of human cognition-inspired evaluation pipeline for voice understanding large language models, LLM Voice, design…
When Large Language Models Meet Speech: A Survey on Integration Approaches
Zhengdong Yang, Shuichiro Shimizu, Yahan Yu +1
Recent advancements in large language models (LLMs) have spurred interest in expanding their application beyond text-based tasks. A large number of studies have explored integratin…
MedKGent: A Large Language Model Agent Framework for Constructing Temporally Evolving Medical Knowledge Graph
Duzhen Zhang, Zixiao Wang, Zhong-Zhi Li +10
The rapid expansion of medical literature challenges the scalable structuring of domain knowledge. Knowledge Graphs (KGs) offer a solution, yet current construction methods lack ge…