activity
20242026
collaborators

7 papers

eess.AS2026

VAPO: End-to-end Slide-Enhanced Speech Recognition with Omni-modal Large Language Models

Rui Hu, Delai Qiu, Yining Wang +2

Omni-modal large language models (OLLMs) offer a promising end-to-end solution for slide-enhanced speech recognition due to their inherent multimodal capabilities. However, we foun…

cs.CV2025

Investigating and Enhancing Vision-Audio Capability in Omnimodal Large Language Models

Rui Hu, Delai Qiu, Shuyu Wei +4

Omnimodal Large Language Models (OLLMs) have shown significant progress in integrating vision and text, but still struggle with integrating vision and audio, often exhibiting subop…

cs.CL2025

ASP2LJ : An Adversarial Self-Play Laywer Augmented Legal Judgment Framework

Ao Chang, Tong Zhou, Yubo Chen +4

Legal Judgment Prediction (LJP) aims to predict judicial outcomes, including relevant legal charge, terms, and fines, which is a crucial process in Large Language Model(LLM). Howev…

cs.CL2025

Semantic Pivots Enable Cross-Lingual Transfer in Large Language Models

Kaiyu He, Tong Zhou, Yubo Chen +4

Large language models (LLMs) demonstrate remarkable ability in cross-lingual tasks. Understanding how LLMs acquire this ability is crucial for their interpretability. To quantify t…

cs.CL2025

Transparentize the Internal and External Knowledge Utilization in LLMs with Trustworthy Citation

Jiajun Shen, Tong Zhou, Yubo Chen +4

While hallucinations of large language models could been alleviated through retrieval-augmented generation and citation generation, how the model utilizes internal knowledge is sti…

cs.CL2025

From Instance Training to Instruction Learning: Task Adapters Generation from Instructions

Huanxuan Liao, Shizhu He, Yao Xu +5

Large language models (LLMs) have acquired the ability to solve general tasks by utilizing instruction finetuning (IFT). However, IFT still relies heavily on instance training of e…