◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Matthew Stallone

3 papers

No researched profile yet.

papers

Publications (3)

cs.LG2024

Diversity Measurement and Subset Selection for Instruction Tuning Datasets

Peiqi Wang, Yikang Shen, Zhen Guo +4

We aim to select data subsets for the fine-tuning of large language models to more effectively follow instructions. Prior work has emphasized the importance of diversity in dataset…

cs.LG2024

Granite-Function Calling Model: Introducing Function Calling Abilities via Multi-task Learning of Granular Tasks

Ibrahim Abdelaziz, Kinjal Basu, Mayank Agarwal +23

Large language models (LLMs) have recently shown tremendous promise in serving as the backbone to agentic systems, as demonstrated by their performance in multi-faceted, challengin…

cs.CL2024

Power Scheduler: A Batch Size and Token Number Agnostic Learning Rate Scheduler

Yikang Shen, Matthew Stallone, Mayank Mishra +6

Finding the optimal learning rate for language model pretraining is a challenging task. This is not only because there is a complicated correlation between learning rate, batch siz…

◍wovepaper

A living map of arXiv — papers, researchers, institutions.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Sign in
  • Library
  • Chat
Data
  • arXiv.org
  • Latest RSS
Metadata from arXiv.org · Not affiliated with arXiv