TinyStories-1MIndonesian Fine-tuning.
An experiment in adapting a small language model to Indonesian text. This project explores the fine-tuning process, training behavior, and the challenges of obtaining useful language generation from a compact model.
Project overview
A small model, a practical learning experiment.
The objective is to explore how an existing language model responds to Indonesian-language fine-tuning, rather than claim a production-quality Indonesian language model.
01 / Project overview
Exploring Indonesian language modeling
A hands-on fine-tuning experiment focused on understanding the process and its limitations.
TinyStories-1M serves as the starting model, while an Indonesian-language dataset is used to explore adapting its learned representations and next-token prediction behavior.
The experiment is intentionally presented with its limitations. Fine-tuning a model on another language does not automatically guarantee fluent, coherent, or accurate generation.
Learning objective
Understand the practical workflow of fine-tuning a small language model.
Language adaptation
Experiment with Indonesian text using Corpus-Indonesia.
Model evaluation
Inspect training loss and perplexity to understand prediction quality.
Experimental mindset
Identify limitations and areas for further training and optimization.
02 / Training setup
The foundation of the experiment
The model, dataset, and hardware form the basis of this fine-tuning run.
Base model
TinyStories-1M
A small language model used as the starting checkpoint.
Training dataset
Corpus-Indonesia
Indonesian-language data from Lyon28.
Explore datasetHardware
NVIDIA T4 GPU
Fine-tuning performed on a T4 GPU.
Training time
> 3 hours
Total reported training duration.
Training dataset: Lyon28/Corpus-Indonesia
An Indonesian-language dataset used for the fine-tuning experiment. Dataset composition, filtering, and data quality can influence the model's final behavior.
03 / Training results
Inspecting loss and perplexity
The reported metrics show that this experiment has substantial room for improvement. These values are presented as reported training results, not as evidence of production-quality language generation.
Lowest listed training loss
5.0924
The lowest loss among the five reported results.
Lowest listed perplexity
162.78
The lowest perplexity among the five reported results.
Training duration
3+ hours
Reported training time using an NVIDIA T4 GPU.
Reported training loss
Relative bar lengths are scaled to the highest listed loss.
These entries are the reported metrics in the model description. Their ordering does not establish that each entry corresponds to a chronological training checkpoint.
Loss and perplexity
Training loss measures how well the model predicts its training targets. Perplexity expresses the uncertainty of next-token prediction on an exponential scale.
Result 1
162.775409
Result 2
302.158057
Result 3
18,700.40634
Result 4
113,509.674623
Result 5
113,546.630401
Results indicate that more work is needed
The reported metrics are not sufficient to demonstrate fluent or reliable Indonesian text generation. A proper evaluation should also include held-out validation data, sample generations, and checks for overfitting and data quality.
04 / Limitations
An experiment, not a finished language model
The project is useful as a learning exercise, but its current reported performance limits its practical applications.
Limitation 01
High perplexity
The reported checkpoints have perplexity ranging from approximately 162.78 to 113,546.63, indicating substantial uncertainty in next-token predictions.
Limitation 02
Training quality
The reported training loss reaches approximately 11.64 in the later listed results. Further investigation is needed to understand the training behavior.
Limitation 03
Experimental only
The model is intended for learning and exploration, not production applications or tasks that require reliable text generation.
What could improve the next experiment?
Investigate learning-rate settings, the number of training epochs, dataset preprocessing, sequence length, and validation performance. Compare generated samples before and after fine-tuning to determine whether the model actually improves on the intended task.
05 / Intended use
What this project is useful for
The value of this experiment is the learning process: understanding fine-tuning, interpreting metrics, and identifying what still needs improvement.
Suitable for exploration
- Learning the full fine-tuning workflow.
- Exploring Indonesian-language model adaptation.
- Inspecting loss and perplexity.
- Experimenting with hyperparameters and datasets.
Not recommended for
- Production text-generation applications.
- Tasks that require high factual accuracy.
- Critical or professional decision-making.
- Applications that require reliable Indonesian generation.
Explore the model
Every experiment is a step toward understanding.
This fine-tuning run is an early exploration of Indonesian language modeling. Its limitations are part of the learning process and provide direction for future experiments.