Learning project · Sole Author · Jun 2025 — Aug 2026
Signature Verification
What the evaluation was hiding
Study project, never deployed
What mattered
Splits are writer-disjoint, so the reported number is not memorised handwriting.
Context
A Siamese ResNet18 that decides whether two signatures came from the same hand.
ResNet18 with conv1 swapped to a single channel, the classifier head stripped,
embeddings L2-normalised, and the decision made on cosine similarity between
them. Trained with CosineEmbeddingLoss on CEDAR: 55 writers, 24 genuine and
24 skilled forgeries each.
I built it in June 2025 and came back a year later to write it up. The write-up turned into a rebuild: writing it up exposed that the old evaluation was overstating the model.
The evaluation was scoring on training data
eval_roc.py loaded signature_pairs.npz and reported an AUC. train.py
loaded the same file. Same 550 pairs, no split — not by pair, not by writer.
The trap underneath: the only saved checkpoint had been trained on all 55 writers, so there was no writer left to test on. Getting an honest number meant retraining from scratch.
Splitting by writer
For a verification system, a random pair split still answers the wrong question. It measures recall of signatures the encoder has already memorised. The real question is whether it works on someone it has never seen, so the split is by writer: 31 train / 8 validation / 16 test, seeded, with an assertion that no writer appears in two groups.
Pair construction matters as much. Negatives are the writer's own skilled forgeries, never another writer's genuine signature — telling two different people apart is nearly free, and would inflate every number.
On those 16 unseen writers the original recipe scores AUC 0.807, EER 27.3%.
Three changes, and where they stopped helping
AUC 0.876, EER 22.9% after:
- 24 pairs per writer instead of 5. The original drew 5 positives from the 276 available combinations, leaving free data on the floor.
- Early stopping on validation writers, selecting on val AUC. The original ran a fixed 20 epochs and hit 100% training accuracy at epoch 9, then spent eleven more epochs memorising, with no validation set that could have noticed.
- Small affine augmentation applied independently to each half of the pair — augmenting both identically leaves the relationship between them unchanged and teaches nothing. Four degrees of rotation, not fifteen: slant is part of what separates a writer from an imitator, and augmenting it away discards real evidence.
Also inverted the images so ink is 1 and paper 0. Preprocessing left ink at 0 on a field of 1s, which meant affine padding (filled with 0) read as ink, and most of the input was constant background.
Then it stops. Validation AUC peaks at epoch one or two and falls while training accuracy climbs to 1.0. The model memorises 24 fixed images per writer almost immediately, and more pairs cannot fix that because every pair recombines the same images. The ceiling is writer count or a triplet-shaped loss, not more epochs.
The explanation was differentiating the wrong scalar
The original Grad-CAM backpropagated torch.ones_like(embedding), the gradient
of the sum of the embedding dimensions. That produces a map of what excites
the encoder, which is not a question anyone asked. The decision is a similarity
between two embeddings, so that is what has to be differentiated:
sim = F.cosine_similarity(e1, e2).squeeze()
sim.backward()
Because both halves pass through the same layer, the forward hook fires twice per explanation and the backward hooks fire in reverse order, so activations and gradients are keyed by call index and paired back up deliberately rather than by hoping the order matches.
It also hooked feature_extractor[3], the maxpool right after the first
convolution, whose heatmap mostly retraces ink. Moved to layer4, and the
5×7 map is upsampled bilinearly rather than by repetition, which reads as a
rendering bug.
What I would not change
The three-way accept / flag / reject decision was right, even though the constants shipped with it were invented rather than fitted — at the hardcoded 0.90 accept cut, 54% of genuine signatures fail to auto-accept on unseen writers. A verification system forced to answer on every input is the wrong shape; the useful output is "a human needs to look at this", and this one at least knew that.
Screens

