Experimental project

TinyStories-1MIndonesian Fine-tuning.

An experiment in adapting a small language model to Indonesian text. This project explores the fine-tuning process, training behavior, and the challenges of obtaining useful language generation from a compact model.

1M-parameter model familyIndonesian languageFull fine-tuning
Not production-ready

Project overview

A small model, a practical learning experiment.

The objective is to explore how an existing language model responds to Indonesian-language fine-tuning, rather than claim a production-quality Indonesian language model.

Base modelTinyStories-1M
DatasetCorpus-Indonesia
TrainingFull fine-tuning
View model card on Hugging Face

01 / Project overview

Exploring Indonesian language modeling

A hands-on fine-tuning experiment focused on understanding the process and its limitations.

TinyStories-1M serves as the starting model, while an Indonesian-language dataset is used to explore adapting its learned representations and next-token prediction behavior.

The experiment is intentionally presented with its limitations. Fine-tuning a model on another language does not automatically guarantee fluent, coherent, or accurate generation.

Learning objective

Understand the practical workflow of fine-tuning a small language model.

Language adaptation

Experiment with Indonesian text using Corpus-Indonesia.

Model evaluation

Inspect training loss and perplexity to understand prediction quality.

Experimental mindset

Identify limitations and areas for further training and optimization.

02 / Training setup

The foundation of the experiment

The model, dataset, and hardware form the basis of this fine-tuning run.

Base model

TinyStories-1M

A small language model used as the starting checkpoint.

Training dataset

Corpus-Indonesia

Indonesian-language data from Lyon28.

Explore dataset

Hardware

NVIDIA T4 GPU

Fine-tuning performed on a T4 GPU.

Training time

> 3 hours

Total reported training duration.

Training dataset: Lyon28/Corpus-Indonesia

An Indonesian-language dataset used for the fine-tuning experiment. Dataset composition, filtering, and data quality can influence the model's final behavior.

Dataset card

03 / Training results

Inspecting loss and perplexity

The reported metrics show that this experiment has substantial room for improvement. These values are presented as reported training results, not as evidence of production-quality language generation.

Lowest listed training loss

5.0924

The lowest loss among the five reported results.

Lowest listed perplexity

162.78

The lowest perplexity among the five reported results.

Training duration

3+ hours

Reported training time using an NVIDIA T4 GPU.

Reported training loss

Relative bar lengths are scaled to the highest listed loss.

Result 15.092371
Result 25.710950
Result 39.836301
Result 411.639643
Result 511.639969

These entries are the reported metrics in the model description. Their ordering does not establish that each entry corresponds to a chronological training checkpoint.

Loss and perplexity

Training loss measures how well the model predicts its training targets. Perplexity expresses the uncertainty of next-token prediction on an exponential scale.

Result 1

162.775409

Perplexity

Result 2

302.158057

Perplexity

Result 3

18,700.40634

Perplexity

Result 4

113,509.674623

Perplexity

Result 5

113,546.630401

Perplexity

Results indicate that more work is needed

The reported metrics are not sufficient to demonstrate fluent or reliable Indonesian text generation. A proper evaluation should also include held-out validation data, sample generations, and checks for overfitting and data quality.

04 / Limitations

An experiment, not a finished language model

The project is useful as a learning exercise, but its current reported performance limits its practical applications.

Limitation 01

High perplexity

The reported checkpoints have perplexity ranging from approximately 162.78 to 113,546.63, indicating substantial uncertainty in next-token predictions.

Limitation 02

Training quality

The reported training loss reaches approximately 11.64 in the later listed results. Further investigation is needed to understand the training behavior.

Limitation 03

Experimental only

The model is intended for learning and exploration, not production applications or tasks that require reliable text generation.

What could improve the next experiment?

Investigate learning-rate settings, the number of training epochs, dataset preprocessing, sequence length, and validation performance. Compare generated samples before and after fine-tuning to determine whether the model actually improves on the intended task.

05 / Intended use

What this project is useful for

The value of this experiment is the learning process: understanding fine-tuning, interpreting metrics, and identifying what still needs improvement.

Suitable for exploration

  • Learning the full fine-tuning workflow.
  • Exploring Indonesian-language model adaptation.
  • Inspecting loss and perplexity.
  • Experimenting with hyperparameters and datasets.

Not recommended for

  • Production text-generation applications.
  • Tasks that require high factual accuracy.
  • Critical or professional decision-making.
  • Applications that require reliable Indonesian generation.

Explore the model

Every experiment is a step toward understanding.

This fine-tuning run is an early exploration of Indonesian language modeling. Its limitations are part of the learning process and provide direction for future experiments.

Open Hugging Face
PythonPyTorchNLPLanguage ModelingFine-tuningHugging Face