activity
20242026
collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL2025

Beyond Global Emotion: Fine-Grained Emotional Speech Synthesis with Dynamic Word-Level Modulation

Sirui Wang, Andong Chen, Tiejun Zhao

Emotional text-to-speech (E-TTS) is central to creating natural and trustworthy human-computer interaction. Existing systems typically rely on sentence-level control through predef…

cs.CL2025

Growing Through Experience: Scaling Episodic Grounding in Language Models

Chunhui Zhang, Sirui, Wang +3

Language models (LMs) require robust episodic grounding-the capacity to learn from and apply past experiences-to excel at physical planning tasks. Current episodic grounding approa…

cs.CL2025

Enhancing LLMs via High-Knowledge Data Selection

Feiyu Duan, Xuemiao Zhang, Sirui Wang +4

The performance of Large Language Models (LLMs) is intrinsically linked to the quality of its training data. Although several studies have proposed methods for high-quality data se…

cs.CL2025

FRAME: Boosting LLMs with A Four-Quadrant Multi-Stage Pretraining Strategy

Xuemiao Zhang, Feiyu Duan, Liangyu Xu +5

Large language models (LLMs) have significantly advanced human language understanding and generation, with pretraining data quality and organization being crucial to their performa…

cs.CL2025

FIRE: Flexible Integration of Data Quality Ratings for Effective Pre-Training

Liangyu Xu, Xuemiao Zhang, Feiyu Duan +4

Selecting high-quality data can improve the pretraining efficiency of large language models (LLMs). Existing methods generally rely on heuristic techniques or single quality signal…

cs.CL2025

Preference Curriculum: LLMs Should Always Be Pretrained on Their Preferred Data

Xuemiao Zhang, Liangyu Xu, Feiyu Duan +5

Large language models (LLMs) generally utilize a consistent data distribution throughout the pretraining process. However, as the model's capability improves, it is intuitive that…