collaborators

5 papers

eess.AS2025

PromptEVC: Controllable Emotional Voice Conversion with Natural Language Prompts

Tianhua Qi, Shiyan Wang, Cheng Lu +4

Controllable emotional voice conversion (EVC) aims to manipulate emotional expressions to increase the diversity of synthesized speech. Existing methods typically rely on predefine…

cs.CV2025

Multimodal Machine Translation with Visual Scene Graph Pruning

Chenyu Lu, Shiliang Sun, Jing Zhao +3

Multimodal machine translation (MMT) seeks to address the challenges posed by linguistic polysemy and ambiguity in translation tasks by incorporating visual information. A key bott…

cs.LG2025

Evaluating Menu OCR and Translation: A Benchmark for Aligning Human and Automated Evaluations in Large Vision-Language Models

Zhanglin Wu, Tengfei Song, Ning Xie +8

The rapid advancement of large vision-language models (LVLMs) has significantly propelled applications in document understanding, particularly in optical character recognition (OCR…

cs.LG2025

Emotion Knowledge Enhancement for Vision Large Language Models: A Self-Verification Approach for High-Quality Emotion Instruction Data Generation

Feifan Wang, Tengfei Song, Minggui He +5

Facial emotion perception in the vision large language model (VLLM) is crucial for achieving natural human-machine interaction. However, creating high-quality annotations for both…

cs.CV2025

DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model

Zhanglin Wu, Tengfei Song, Ning Xie +6

This paper presents the technical solution proposed by Huawei Translation Service Center (HW-TSC) for the "End-to-End Document Image Machine Translation for Complex Layouts" compet…