collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2024

Adapting Safe-for-Work Classifier for Malaysian Language Text: Enhancing Alignment in LLM-Ops Framework

Aisyah Razak, Ariff Nazhan, Kamarul Adha +3

As large language models (LLMs) become increasingly integrated into operational workflows (LLM-Ops), there is a pressing need for effective guardrails to ensure safe and aligned in…

cs.CL2024

MMMModal -- Multi-Images Multi-Audio Multi-turn Multi-Modal

Husein Zolkepli, Aisyah Razak, Kamarul Adha +1

Our contribution introduces a groundbreaking multimodal large language model designed to comprehend multi-images, multi-audio, and multi-images-multi-audio within a single multitur…

cs.CL2024

Multi-Lingual Malaysian Embedding: Leveraging Large Language Models for Semantic Representations

Husein Zolkepli, Aisyah Razak, Kamarul Adha +1

In this work, we present a comprehensive exploration of finetuning Malaysian language models, specifically Llama2 and Mistral, on embedding tasks involving negative and positive pa…

cs.CL2024

Large Malaysian Language Model Based on Mistral for Enhanced Local Language Understanding

Husein Zolkepli, Aisyah Razak, Kamarul Adha +1

In this paper, we present significant advancements in the pretraining of Mistral 7B, a large-scale language model, using a dataset of 32.6 GB, equivalent to 1.1 billion tokens. We…

cs.CL2024

MaLLaM -- Malaysia Large Language Model

Husein Zolkepli, Aisyah Razak, Kamarul Adha +1

Addressing the gap in Large Language Model pretrained from scratch with Malaysian context, We trained models with 1.1 billion, 3 billion, and 5 billion parameters on a substantial…