Under the hood of EasyOCR

How EasyOCRactually works.

An end-to-end walkthrough of EasyOCR internals, from CRAFT text detection and bounding-box postprocessing to VGG-based recognition and CTC decoding.

CRAFT detectorVGG recognizerCTC decoding

Final inference

Text detection and recognition

Output
EasyOCR final inference output showing detected text boxes and confidence scores
Final prediction: bounding boxes, recognized text, and confidence scores.

The motivation

Understand the implementation, not just the API.

EasyOCR makes text extraction straightforward through a single inference interface. This project opens up that interface to understand how image pixels eventually become recognized text.

Each notebook isolates a component of the pipeline, making intermediate tensors, score maps, thresholds, box matching, and decoder behavior easier to inspect and experiment with.

Browse the notebooks

Detection internals

Explore CRAFT score maps, thresholding, connected components, and text-box generation.

Threshold experiments

See how changing low_text affects active pixels and candidate text regions.

Recognition internals

Inspect the VGG-based recognizer, feature sequences, and character predictions.

Decoder comparison

Study greedy search, beam search, word beam search, and repeated-character handling.

01 / Pipeline overview

From pixels to recognized text

The inference process is a sequence of detection, geometric postprocessing, region preparation, recognition, and decoding.

01

Image preprocessing

Resize the input image, preserve its aspect ratio, normalize pixel values, and convert it into a tensor.

02

CRAFT detection

Predict character-region and affinity scores to identify text regions and connections between neighboring characters.

03

Postprocessing

Apply thresholds, extract connected components, estimate bounding boxes, and group nearby detections.

04

Region cropping

Crop detected text regions and transform them into inputs that the recognition model can process.

05

Text recognition

Pass each cropped region through the VGG-based feature extractor and sequence prediction layers.

06

CTC decoding

Convert timestep-level character predictions into readable text using greedy search, beam search, or word beam search.

Detection and recognition are separate tasks

CRAFT determines where text is located. The recognition model determines what that text says.

02 / Text detection

CRAFT: detecting text at the pixel level

CRAFT predicts two score maps that help identify character regions and the affinities connecting characters into words.

CRAFT text score and link score heatmaps
CRAFT inference outputs: text-region score and character-link score.
T

Text score

Estimates how likely each pixel belongs to a character region. High-scoring pixels form the foundation for detecting text components.

L

Link score

Estimates the affinity between neighboring characters. These connections help the postprocessing stage group characters that belong to the same word.

Key insight

The detector does not directly output the final recognized text. Its score maps must first be converted into geometric text regions.

03 / Threshold analysis

How does low_text change detection?

Thresholding converts continuous score maps into binary regions. Experimenting with low_text reveals how sensitive candidate text regions are to this parameter.

Parameter study

Low-text threshold vs. active pixels

The low-text threshold determines which pixels are eligible to participate in candidate text regions. Lowering the threshold can retain weaker responses, while raising it can remove faint regions.

text_score > low_text

Conceptual thresholding rule. Actual box generation also depends on link scores, connected components, and subsequent filtering.

Graph showing the effect of the low text threshold on the number of active pixels
Experiment: how the low_text threshold affects the number of active pixels.
Visual comparison of text detection under different low text threshold values
Side-by-side comparison of detection masks produced with different low_text values.
1

Low threshold

A lower threshold can activate more pixels, revealing weaker text regions but potentially introducing noise.

2

Connected regions

Active pixels are converted into connected components that become candidates for text boxes.

3

Final boxes

Candidate regions are filtered and grouped according to their geometry and neighboring text relationships.

Note: the exact pixel counts and threshold values should be interpreted from the plotted experiment, rather than assumed from the parameter name alone.

04 / Detection postprocessing

From score maps to bounding boxes

After thresholding, EasyOCR converts candidate pixels into connected components, estimates text geometry, and matches or groups regions before recognition.

Visualization of matching and grouping candidate text boxes
Matching boxes: analyzing relationships between neighboring candidate regions.

What happens here?

  • Extract connected components from the thresholded score maps.
  • Estimate bounding boxes or polygons around candidate regions.
  • Match neighboring components when they likely belong together.
  • Group text boxes into suitable recognition regions.

Why are boxes matched?

Text does not always appear as one clean connected region. Characters may produce separate components, and neighboring text may need to be grouped. Postprocessing uses geometry and connectivity to form useful regions for recognition.

Not every candidate should be merged. Incorrect grouping can combine separate words, while insufficient grouping can fragment a single word into multiple regions.

The output is still geometry

At this stage, the pipeline has candidate text regions and their coordinates. The actual characters are predicted by the recognizer in the next stage.

Bounding boxesCropped text regions

05 / Text recognition

Each detected region becomes a sequence

Once text regions are prepared, the recognizer predicts a sequence of character probabilities for each crop.

01

Crop

Extract each detected region from the original image.

02

Normalize

Resize and normalize the crop to match recognizer input requirements.

03

Predict

Generate a sequence of character scores across feature timesteps.

English Generation 2 recognizer

This walkthrough inspects the English Generation 2 recognition model, including its VGG-based feature extractor, sequence modeling, prediction layer, and CTC output probabilities.

VGG feature extractorSequence modelingPrediction layerCharacter probabilities

06 / Sequence decoding

CTC turns predictions into text

The recognizer emits character probabilities over timesteps. CTC decoding converts those predictions into a readable character sequence.

Model output

Timestep predictions

A sequence of probability distributions over the character vocabulary, including the CTC blank token.

Decoder output

Recognized characters

Collapse repeated labels and remove blank tokens to obtain the decoded sequence.

Three decoding strategies explored

Greedy search

Select the most likely character at each timestep before collapsing the sequence.

Beam search

Maintain multiple candidate paths to explore alternatives beyond a single greedy path.

Word beam search

Use word-level constraints to guide decoding toward plausible word sequences.

Why inspect the decoder?

Repeated characters, blank-token removal, confidence calculations, and beam-search behavior all affect how timestep predictions become the final text. The notebooks also investigate CTC-related decoding behavior.

07 / Final result

The complete inference output

The final result combines the geometry from text detection with the recognized text and its associated confidence score.

Final EasyOCR prediction with text bounding boxes, recognized text, and confidence scores
End-to-end output from the inspected EasyOCR pipeline.

Bounding boxes

Coordinates define the detected text regions in the image.

Recognized text

The recognition and decoding stages produce readable strings.

Confidence score

A confidence value accompanies each recognition result.

Explore the implementation

Learn EasyOCR from the inside out.

Explore the notebooks covering preprocessing, CRAFT inference, box postprocessing, recognition, and CTC decoding. Each experiment focuses on making one part of the pipeline observable and understandable.

View on GitHub
PythonPyTorchCRAFTComputer VisionOCRCTC