PyTorchSwin TransformerBARTMLflowVertex AIHugging Face

Invoice Documentto Text

An end-to-end deep learning system that converts invoice images directly into structured JSON using a custom Swin Transformer + BART encoder-decoder architecture.

Model Demo

From invoice image to structured JSON

The model processes the document as a single visual input and generates the structured invoice representation directly.

Input

Synthetic invoice image

IMAGE
Synthetic invoice used for model inference

Output

Generated structured data

JSON
{
  "invoice_date": "1 February 2025",
  "invoice_no": "INV-2025-30496",
  "subtotal": "84750000",
  "from_to": {
    "name": "PT Solusi Nusantara",
    "city": "Meulaboh, Indonesia",
    "email": "billing@solusinusantara.com"
  },
  "bill_to": {
    "name": "PT Maju Teknologi",
    "city": "Bontang, Indonesia",
    "email": "finance@majuteknologi.com"
  },
  "items": [
    {
      "qty": "1",
      "description": "DevOps Services",
      "price": "3800000",
      "total": "3800000"
    },
    {
      "qty": "9",
      "description": "UI/UX Design",
      "price": "2900000",
      "total": "26100000"
    },
    {
      "qty": "8",
      "description": "UI/UX Design",
      "price": "4000000",
      "total": "32000000"
    },
    {
      "qty": "1",
      "description": "Cloud Infrastructure",
      "price": "1250000",
      "total": "1250000"
    },
    {
      "qty": "7",
      "description": "AI Development",
      "price": "2400000",
      "total": "16800000"
    },
    {
      "qty": "3",
      "description": "Cloud Infrastructure",
      "price": "1600000",
      "total": "4800000"
    }
  ],
  "payment_information": {
    "bank": "Bank Negara Indonesia",
    "account_number": "1438989805",
    "account_name": "PT Solusi Nusantara"
  }
}

This example demonstrates the complete image-to-structured-data flow: the invoice image is processed by the trained model and decoded into structured invoice information.

Architecture

A custom vision-to-sequence pipeline

Instead of combining a separate OCR engine with an LLM, the project explores an end-to-end document understanding architecture.

01

Invoice Image

Visual document input

02

Preprocessing

Resize & normalize

03

Swin Encoder

Visual feature extraction

04

BART Decoder

Autoregressive generation

05

Structured JSON

Parsed invoice data

Swin Transformer Encoder

The visual encoder transforms the invoice image into visual representations containing information about text regions, layout, tables, headers, rows, and spatial relationships.

BART Decoder

The decoder receives visual features through cross-attention and generates the target representation token by token until the structured invoice sequence is complete.

Dataset

Synthetic invoice data

The model is trained on a synthetic dataset designed specifically for controlled invoice document understanding experiments.

1,000

Total synthetic invoices

900 / 100

Train / test samples

Image + JSON

Paired training data

synth-invoice

Synthetic invoice images paired with structured ground truth, used to train and evaluate the image-to-text pipeline.

Open Dataset

Training & MLOps

From local experimentation to GPU training

The project goes beyond model architecture and covers the surrounding workflow required to train, track, package, and publish the model.

PyTorch

Model implementation, dataset loading, training loop, checkpointing, and inference.

MLflow

Tracks experiments, training configuration, metrics, checkpoints, and model artifacts.

Vertex AI

Provides GPU-based training workflows for larger model training runs and deployment-oriented experiments.

Dataset

Preprocessing

Model Training

MLflow

Hugging Face

Result

Image in. Structured data out.

The project demonstrates how an invoice document can be processed end-to-end by a custom vision-to-sequence model without relying on a separate OCR engine followed by an LLM.

Main Code

Training & inference

The core implementation uses Hugging Face Transformers and Donut-style encoder-decoder components for training the invoice document understanding model and running inference.

Dataset

synth-invoice

Base Model

donut-base

Training

20 epochs ยท FP16

python
from datasets import load_dataset
from transformers import DonutProcessor, DonutModel

dataset = load_dataset("ridwanFatur98/synth-invoice")

processor = DonutProcessor.from_pretrained(
    "naver-clova-ix/donut-base"
)

model = DonutModel.from_pretrained(
    "naver-clova-ix/donut-base"
)

new_special_tokens = [
    "</s_qty>",
    "</s_bill_to>",
    "</s_account_name>",
    "<s_total>",
    "<s_description>",
    "</s_account_number>",
    "<s_invoice_date>",
    "<s_items>",
    "<s_account_name>",
    "<s_invoice_no>",
    "<s_note>",
    "</s_payment_information>",
    "<s_bill_to>",
    "</s_name>",
    "<s_address>",
    "</s_price>",
    "<parsing>",
    "<s_account_number>",
    "</s_note>",
    "</s_address>",
    "<s_phone>",
    "<s_qty>",
    "</s_subtotal>",
    "</s_city>",
    "<s_from_to>",
    "</s_items>",
    "</s_invoice_no>",
    "</s_bank>",
    "</s_email>",
    "<s_subtotal>",
    "</s_total>",
    "</s_phone>",
    "<s_name>",
    "<s_city>",
    "<s_price>",
    "<s_email>",
    "<s_payment_information>",
    "<s_bank>",
    "</s_from_to>",
    "</s_invoice_date>",
    "</s_description>",
]

# Arguments for training
from transformers import Seq2SeqTrainingArguments, Seq2SeqTrainer

training_args = Seq2SeqTrainingArguments(
    output_dir=save_model_name,
    num_train_epochs=20,
    max_steps=-1,

    learning_rate=2e-5,
    per_device_train_batch_size=1,
    per_device_eval_batch_size=1,
    weight_decay=0.01,

    fp16=True,
    logging_steps=100,

    eval_strategy="no",
    save_strategy="epoch",
    save_total_limit=1,

    predict_with_generate=True,
    report_to="tensorboard",
    push_to_hub=False,
    remove_unused_columns=False,
)

trainer = Seq2SeqTrainer(
    model=model,
    args=training_args,
    data_collator=prepare_data,
    train_dataset=proc_dataset["train"],
    eval_dataset=proc_dataset["test"],
)

trainer.train()