All posts

Machine learning

Recognising handwritten letters with a CNN and OpenCV, then making it production ready

A hands-on walkthrough of my A to Z handwriting recogniser in Python: data loading, OpenCV preprocessing, a Keras CNN, honest evaluation, and how I would version, track and serve it today.

Published 11 min read

A few years ago I put together a small project that reads a photo of a single handwritten capital letter and tells you which of the 26 it is. The code lives in AshabaJasper/Handwriting-Recognition-in-Python (opens in a new tab). It started from the DataFlair handwritten character recognition walkthrough (the OpenCV window title still says so), and it is a single script: load a CSV, train a convolutional neural network, then classify an image from disk.

In this post I rebuild that script step by step with current libraries, explain why each step exists, point out what the original got wrong, and finish with how I would take it from a script to something you can version, track and serve.

What you need

The original README pins NumPy 1.16, OpenCV 3.4, Keras 2.3 and TensorFlow 2.0. Those versions are long gone, so the excerpts below target TensorFlow 2 with the bundled Keras, a current NumPy, OpenCV (opencv-python-headless is enough), pandas and scikit-learn.

The data is the A to Z handwritten alphabets dataset on Kaggle (opens in a new tab). It is one CSV where each row is a letter: the first column is the label (0 for A up to 25 for Z) and the remaining 784 columns are the pixels of a 28 by 28 greyscale image, white ink on a black background. Download it and put it in a data/ folder.

Step 1: load and split the data

The original script reads the CSV, drops the label column (whose header is literally "0") to get the pixels, and splits 80/20 into train and test. I keep that, but add three things: a stratified split so every letter keeps its share in both halves, a fixed random seed so the split is reproducible, and scaling the pixels from 0 to 255 down to 0 to 1, which the original skipped.

File: data.py
import numpy as np
import pandas as pd
from sklearn.model_selection import train_test_split
 
LETTERS = [chr(ord("A") + i) for i in range(26)]
CSV_PATH = "data/A_Z Handwritten Data.csv"
 
 
def load_splits(seed: int = 42):
    data = pd.read_csv(CSV_PATH).astype("float32")
    X = data.drop("0", axis=1).to_numpy()
    y = data["0"].to_numpy().astype("int64")
 
    train_x, test_x, train_y, test_y = train_test_split(
        X, y, test_size=0.2, stratify=y, random_state=seed
    )
    # Pixels arrive flat; the CNN wants (height, width, channels).
    train_x = train_x.reshape(-1, 28, 28, 1) / 255.0
    test_x = test_x.reshape(-1, 28, 28, 1) / 255.0
    return train_x, test_x, train_y, test_y

Scaling is cheap and usually makes training steadier, because the first layer no longer sees very large activations.

Look at the data before you model it

The original plotted a horizontal bar chart of how many samples each letter has, and a 3 by 3 grid of random images. Both are worth keeping. The bar chart tells you straight away that the classes are not balanced, which is the reason I later look at per-class metrics rather than a single accuracy figure.

File: explore.py
import matplotlib.pyplot as plt
import numpy as np
 
from data import LETTERS, load_splits
 
train_x, test_x, train_y, test_y = load_splits()
 
counts = np.bincount(train_y, minlength=26)
fig, ax = plt.subplots(figsize=(8, 8))
ax.barh(LETTERS, counts)
ax.set_xlabel("Training samples")
ax.set_ylabel("Letter")
plt.show()
 
fig, axes = plt.subplots(3, 3, figsize=(6, 6))
for ax, idx in zip(axes.flat, np.random.default_rng(0).choice(len(train_x), 9)):
    ax.imshow(train_x[idx].squeeze(), cmap="Greys")
    ax.set_title(LETTERS[train_y[idx]])
    ax.axis("off")
plt.show()

The original counted labels with np.int0, which NumPy 2 removed; np.bincount replaces it.

Step 2: the model

The network in the repo is a classic small CNN: three convolution and max pooling blocks that grow from 32 to 64 to 128 filters, then two dense layers and a 26-way softmax. I have kept the architecture exactly as it was, only declaring the input with keras.Input, which current Keras prefers.

File: model.py
import keras
from keras import layers
 
 
def build_model() -> keras.Model:
    model = keras.Sequential([
        keras.Input(shape=(28, 28, 1)),
        layers.Conv2D(32, (3, 3), activation="relu"),
        layers.MaxPool2D(pool_size=(2, 2), strides=2),
        layers.Conv2D(64, (3, 3), activation="relu", padding="same"),
        layers.MaxPool2D(pool_size=(2, 2), strides=2),
        layers.Conv2D(128, (3, 3), activation="relu", padding="valid"),
        layers.MaxPool2D(pool_size=(2, 2), strides=2),
        layers.Flatten(),
        layers.Dense(64, activation="relu"),
        layers.Dense(128, activation="relu"),
        layers.Dense(26, activation="softmax"),
    ])
    model.compile(
        optimizer=keras.optimizers.Adam(learning_rate=0.001),
        loss="sparse_categorical_crossentropy",
        metrics=["accuracy"],
    )
    return model

The convolutions learn local stroke patterns such as curves, corners and crossings. Pooling halves the spatial size after each block, so later filters see a larger part of the letter. By the time we flatten, the image has become a short vector of features, and the dense layers map that to a probability for each letter.

The original one-hot encoded the labels with to_categorical and used categorical_crossentropy. Using sparse_categorical_crossentropy with integer labels is equivalent and saves memory and a step.

Step 3: training

The original used Adam with two callbacks: ReduceLROnPlateau, which cuts the learning rate by a factor of five when validation loss stalls, and EarlyStopping, which stops when it stops improving. Both are sensible. Two things were not.

Two problems in the original training step

First, it trained for a single epoch, so the callbacks never had a chance to do anything. Second, it passed the test set as validation_data, which means the test set was steering the training. Any score measured on it afterwards is optimistic. Below I carve a validation set out of the training data and keep the test set untouched until the very end.

File: train.py
import keras
 
from data import load_splits
from model import build_model
 
train_x, test_x, train_y, test_y = load_splits()
model = build_model()
 
callbacks = [
    keras.callbacks.ReduceLROnPlateau(
        monitor="val_loss", factor=0.2, patience=1, min_lr=0.0001
    ),
    keras.callbacks.EarlyStopping(
        monitor="val_loss", patience=2, restore_best_weights=True
    ),
]
 
history = model.fit(
    train_x,
    train_y,
    epochs=20,
    batch_size=128,
    validation_split=0.1,  # taken from training data, not the test set
    callbacks=callbacks,
)
model.save("models/model_hand.keras")

Twenty epochs is a ceiling, not a target: early stopping ends the run when validation loss stops improving, and restore_best_weights=True keeps the best model rather than the last. I also switched from .h5 to the native .keras format.

Step 4: evaluation

The repo does not include any saved run output, so I cannot quote an accuracy figure for it, and I would rather not invent one. What I can show is how to evaluate properly once you have trained it yourself.

File: evaluate.py
import keras
import numpy as np
from sklearn.metrics import classification_report, confusion_matrix
 
from data import LETTERS, load_splits
 
_, test_x, _, test_y = load_splits()
model = keras.models.load_model("models/model_hand.keras")
 
loss, accuracy = model.evaluate(test_x, test_y, verbose=0)
print(f"Test loss {loss:.4f}, test accuracy {accuracy:.4f}")
 
pred = model.predict(test_x, verbose=0).argmax(axis=1)
print(classification_report(test_y, pred, target_names=LETTERS, digits=3))
 
# The most confused pairs say more than the headline number.
cm = confusion_matrix(test_y, pred)
np.fill_diagonal(cm, 0)
for flat in cm.ravel().argsort()[::-1][:5]:
    true_i, pred_i = divmod(int(flat), 26)
    print(f"{LETTERS[true_i]} read as {LETTERS[pred_i]}: {cm[true_i, pred_i]}")

Because the classes are imbalanced, overall accuracy can look healthy while a rare letter does badly. The per-class precision and recall in the report, and the list of most confused pairs, are where you find the real weaknesses.

There is also a bug in the original worth calling out. Its 3 by 3 "Prediction" grid of test images took the title from the one-hot test labels, not from model.predict. So the grid always showed the right answer, whatever the model thought. If you plot predictions, take them from pred above.

Step 5: reading a real photo with OpenCV

This is the part I enjoyed most. A photo of a letter on paper looks nothing like the training data: it is in colour, the ink is dark on a light background, the lighting is uneven, and the letter is somewhere in a large frame. The original handled this with a Gaussian blur, a conversion to greyscale, an inverted binary threshold at a fixed value of 100, and a resize of the whole frame to 28 by 28.

That works when the letter fills the photo. When it does not, the letter shrinks to a few pixels. Here is the version I would use now, still built from the same OpenCV steps, plus cropping to the ink and centring it:

File: preprocess.py
import cv2
import numpy as np
 
 
def preprocess(gray: np.ndarray) -> np.ndarray:
    """Greyscale photo of one dark letter on light paper -> model input."""
    blurred = cv2.GaussianBlur(gray, (7, 7), 0)
    # Invert so ink is white on black, like the training images. Otsu picks
    # the threshold per image instead of the fixed 100 in the original.
    _, ink = cv2.threshold(
        blurred, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU
    )
    points = cv2.findNonZero(ink)
    if points is None:
        raise ValueError("No ink found in the image")
    x, y, w, h = cv2.boundingRect(points)
    glyph = ink[y : y + h, x : x + w]
 
    # Pad to a square so the letter keeps its proportions, then leave a
    # margin around it inside the 28 by 28 frame.
    side = max(w, h)
    square = np.zeros((side, side), dtype=np.uint8)
    top, left = (side - h) // 2, (side - w) // 2
    square[top : top + h, left : left + w] = glyph
    small = cv2.resize(square, (20, 20), interpolation=cv2.INTER_AREA)
    framed = np.pad(small, 4)
    return framed.astype("float32").reshape(1, 28, 28, 1) / 255.0

The margin of four pixels is my choice to roughly match how letters sit in the training images. Check it against your data by plotting a few preprocessed photos next to a few training samples; if they look different to you, they look different to the model too. Predicting is then one line: LETTERS[model.predict(preprocess(gray)).argmax()], where gray comes from cv2.imread(path, cv2.IMREAD_GRAYSCALE).

Step 6: from script to something you can run in production

The script trains, saves and predicts in one go, with hard-coded Windows paths. That is fine for learning. For anything that other people rely on, I would split it into the files above and add three things: versioned data, tracked runs, and a small service.

Version the data

A model is only reproducible if you know exactly which data trained it. I use DVC (opens in a new tab) for this: dvc add "data/A_Z Handwritten Data.csv" replaces the large file in Git with a small pointer file holding its hash, and dvc push sends the data itself to remote storage such as S3 or a shared drive. Anyone who checks out a commit and runs dvc pull gets the same bytes. Adding the steps to a dvc.yaml pipeline means dvc repro reruns only the stages whose inputs changed.

Track every run

Next, record what each training run used and produced. MLflow is easy to start with and runs locally with no server:

File: train_tracked.py
import hashlib
 
import mlflow
 
 
def file_md5(path: str) -> str:
    digest = hashlib.md5()
    with open(path, "rb") as handle:
        for chunk in iter(lambda: handle.read(1 << 20), b""):
            digest.update(chunk)
    return digest.hexdigest()
 
 
mlflow.set_experiment("handwriting-a-to-z")
with mlflow.start_run():
    mlflow.log_params({
        "data_md5": file_md5("data/A_Z Handwritten Data.csv"),
        "split_seed": 42,
        "learning_rate": 0.001,
        "max_epochs": 20,
    })
    # ... build, fit and evaluate as in train.py and evaluate.py ...
    mlflow.log_metric("test_accuracy", accuracy)
    mlflow.log_artifact("models/model_hand.keras")

Now every model file is linked to a data hash, its settings and its test score, and mlflow ui lets you compare runs side by side. When a new model is better on the held-out test set and no worse on the letters you care about, promote it; otherwise keep the old one.

Serve it behind a small API

For serving, a FastAPI app that loads the model once and accepts an uploaded image is plenty:

File: serve.py
import cv2
import keras
import numpy as np
from fastapi import FastAPI, HTTPException, UploadFile
 
from data import LETTERS
from preprocess import preprocess
 
app = FastAPI()
model = keras.models.load_model("models/model_hand.keras")
 
 
@app.post("/predict")
async def predict(file: UploadFile):
    raw = np.frombuffer(await file.read(), dtype=np.uint8)
    gray = cv2.imdecode(raw, cv2.IMREAD_GRAYSCALE)
    if gray is None:
        raise HTTPException(status_code=400, detail="Not a readable image")
    try:
        probs = model.predict(preprocess(gray), verbose=0)[0]
    except ValueError as err:
        raise HTTPException(status_code=422, detail=str(err))
    top = probs.argsort()[::-1][:3]
    return {"top3": [{"letter": LETTERS[i], "p": float(probs[i])} for i in top]}

Returning the top three with probabilities lets the client decide what to do with an uncertain answer. Run it with uvicorn serve:app and put it in a slim Docker image with pinned dependencies. In a real service I would also limit upload size and log the predicted letter and its probability (not the image) to spot drift.

Test the fragile part

The model is not the part most likely to break; the preprocessing is. A tiny pytest guards it:

File: test_preprocess.py
import cv2
import numpy as np
import pytest
 
from preprocess import preprocess
 
 
def test_letter_becomes_model_input():
    page = np.full((200, 300), 255, dtype=np.uint8)
    cv2.putText(page, "A", (120, 140), cv2.FONT_HERSHEY_SIMPLEX, 3, 0, 8)
    out = preprocess(page)
    assert out.shape == (1, 28, 28, 1)
    assert 0.0 <= out.min() and out.max() <= 1.0
    assert out.max() > 0.5  # the letter survived
 
 
def test_blank_page_is_rejected():
    with pytest.raises(ValueError):
        preprocess(np.full((100, 100), 255, dtype=np.uint8))

What I would improve next

Where the real gains are

The architecture is rarely the bottleneck here. Data augmentation (small rotations, shifts and thickness changes with keras.layers.RandomRotation and friends), class weights or resampling for the rarer letters, and collecting a few hundred real phone photos as an extra test set will tell you far more about real-world performance than another convolution layer.

Beyond that, reading whole words is a natural next step: find each letter with cv2.findContours, sort the boxes left to right, and classify each crop. Joined-up handwriting is a different problem and needs a sequence model.

If you try this, start from the repository (opens in a new tab), split it into the files above, and train it yourself. The numbers you get on your own held-out photos are the ones worth trusting.

Comments

Every comment is read before it appears here. Be kind and stay on topic; your email is never published.

Loading comments...

Leave a comment

Never shown. Only used if I need to reply privately.