Invoice Documentto Text
An end-to-end deep learning system that converts invoice images directly into structured JSON using a custom Swin Transformer + BART encoder-decoder architecture.
Model Demo
From invoice image to structured JSON
The model processes the document as a single visual input and generates the structured invoice representation directly.
Input
Synthetic invoice image

Output
Generated structured data
{
"invoice_date": "1 February 2025",
"invoice_no": "INV-2025-30496",
"subtotal": "84750000",
"from_to": {
"name": "PT Solusi Nusantara",
"city": "Meulaboh, Indonesia",
"email": "billing@solusinusantara.com"
},
"bill_to": {
"name": "PT Maju Teknologi",
"city": "Bontang, Indonesia",
"email": "finance@majuteknologi.com"
},
"items": [
{
"qty": "1",
"description": "DevOps Services",
"price": "3800000",
"total": "3800000"
},
{
"qty": "9",
"description": "UI/UX Design",
"price": "2900000",
"total": "26100000"
},
{
"qty": "8",
"description": "UI/UX Design",
"price": "4000000",
"total": "32000000"
},
{
"qty": "1",
"description": "Cloud Infrastructure",
"price": "1250000",
"total": "1250000"
},
{
"qty": "7",
"description": "AI Development",
"price": "2400000",
"total": "16800000"
},
{
"qty": "3",
"description": "Cloud Infrastructure",
"price": "1600000",
"total": "4800000"
}
],
"payment_information": {
"bank": "Bank Negara Indonesia",
"account_number": "1438989805",
"account_name": "PT Solusi Nusantara"
}
}This example demonstrates the complete image-to-structured-data flow: the invoice image is processed by the trained model and decoded into structured invoice information.
Architecture
A custom vision-to-sequence pipeline
Instead of combining a separate OCR engine with an LLM, the project explores an end-to-end document understanding architecture.
01
Invoice Image
Visual document input
02
Preprocessing
Resize & normalize
03
Swin Encoder
Visual feature extraction
04
BART Decoder
Autoregressive generation
05
Structured JSON
Parsed invoice data
Swin Transformer Encoder
The visual encoder transforms the invoice image into visual representations containing information about text regions, layout, tables, headers, rows, and spatial relationships.
BART Decoder
The decoder receives visual features through cross-attention and generates the target representation token by token until the structured invoice sequence is complete.
Dataset
Synthetic invoice data
The model is trained on a synthetic dataset designed specifically for controlled invoice document understanding experiments.
1,000
Total synthetic invoices
900 / 100
Train / test samples
Image + JSON
Paired training data
synth-invoice
Synthetic invoice images paired with structured ground truth, used to train and evaluate the image-to-text pipeline.
Development Workflow
Built progressively from individual components
The project is organized into notebooks so each stage of the document understanding pipeline can be developed and validated independently.
Training & MLOps
From local experimentation to GPU training
The project goes beyond model architecture and covers the surrounding workflow required to train, track, package, and publish the model.
PyTorch
Model implementation, dataset loading, training loop, checkpointing, and inference.
MLflow
Tracks experiments, training configuration, metrics, checkpoints, and model artifacts.
Vertex AI
Provides GPU-based training workflows for larger model training runs and deployment-oriented experiments.
Dataset
Preprocessing
Model Training
MLflow
Hugging Face
Result
Image in. Structured data out.
The project demonstrates how an invoice document can be processed end-to-end by a custom vision-to-sequence model without relying on a separate OCR engine followed by an LLM.
Main Code
Training & inference
The core implementation uses Hugging Face Transformers and Donut-style encoder-decoder components for training the invoice document understanding model and running inference.
Dataset
synth-invoice
Base Model
donut-base
Training
20 epochs ยท FP16
from datasets import load_dataset
from transformers import DonutProcessor, DonutModel
dataset = load_dataset("ridwanFatur98/synth-invoice")
processor = DonutProcessor.from_pretrained(
"naver-clova-ix/donut-base"
)
model = DonutModel.from_pretrained(
"naver-clova-ix/donut-base"
)
new_special_tokens = [
"</s_qty>",
"</s_bill_to>",
"</s_account_name>",
"<s_total>",
"<s_description>",
"</s_account_number>",
"<s_invoice_date>",
"<s_items>",
"<s_account_name>",
"<s_invoice_no>",
"<s_note>",
"</s_payment_information>",
"<s_bill_to>",
"</s_name>",
"<s_address>",
"</s_price>",
"<parsing>",
"<s_account_number>",
"</s_note>",
"</s_address>",
"<s_phone>",
"<s_qty>",
"</s_subtotal>",
"</s_city>",
"<s_from_to>",
"</s_items>",
"</s_invoice_no>",
"</s_bank>",
"</s_email>",
"<s_subtotal>",
"</s_total>",
"</s_phone>",
"<s_name>",
"<s_city>",
"<s_price>",
"<s_email>",
"<s_payment_information>",
"<s_bank>",
"</s_from_to>",
"</s_invoice_date>",
"</s_description>",
]
# Arguments for training
from transformers import Seq2SeqTrainingArguments, Seq2SeqTrainer
training_args = Seq2SeqTrainingArguments(
output_dir=save_model_name,
num_train_epochs=20,
max_steps=-1,
learning_rate=2e-5,
per_device_train_batch_size=1,
per_device_eval_batch_size=1,
weight_decay=0.01,
fp16=True,
logging_steps=100,
eval_strategy="no",
save_strategy="epoch",
save_total_limit=1,
predict_with_generate=True,
report_to="tensorboard",
push_to_hub=False,
remove_unused_columns=False,
)
trainer = Seq2SeqTrainer(
model=model,
args=training_args,
data_collator=prepare_data,
train_dataset=proc_dataset["train"],
eval_dataset=proc_dataset["test"],
)
trainer.train()