Amir TabatabaeiWork with me
← All work

Learning project · Sole Author · May 2025 — Aug 2026

CIFAR-10 Classifiers

An 86× larger network that lost

Study project, never deployed

PyTorchtorchvisionResNet18scikit-learn

What mattered

The ImageNet stem throws away a 32×32 image before the first residual block; the CIFAR stem keeps it.

System map

IMAGENET/2/2CIFARNO POOL32x32 input7x7 stride 2maxpool /28x83x3 stride 132x32

Context

Two classifiers trained from scratch on CIFAR-10: a hand-written TinyVGG and torchvision's ResNet18. Three epochs each, batch 32, Adam at 1e-3, the same flip-and-crop augmentation.

Both checkpoints were still in the repo a year later, so I could measure them rather than trust the training logs. The result is the reason this is worth a page:

| | Parameters | Test accuracy | |---|---|---| | TinyVGG | 0.13M | 68.30% | | ResNet18 | 11.18M | 66.69% |

Eighty-six times the parameters, and it loses. Same budget, same data, same augmentation — so it is not undertraining.

The stem throws the image away

torchvision's ResNet18 is built for 224×224 ImageNet input. Its stem is a 7×7 stride-2 convolution followed by a 3×3 stride-2 maxpool — a 4× linear reduction before any residual block runs. On a 224×224 image that is sensible. On a 32×32 image it is most of the picture.

Rather than argue it, I measured the tensor:

ResNet18 ImageNet stem: 32x32 -> (8, 8) entering layer1
ResNet18 CIFAR stem:    32x32 -> (32, 32) entering layer1

Eight by eight. Then layer2 through layer4 downsample three more times, to 1×1. Eleven million parameters are operating on an image that no longer exists, which is why a 0.13M-parameter network that kept its resolution beat it.

Two lines, +12.1 points

The standard CIFAR adaptation replaces the stem and drops the maxpool:

m.conv1 = nn.Conv2d(3, 64, kernel_size=3, stride=1, padding=1, bias=False)
m.maxpool = nn.Identity()

I retrained with everything else held fixed — same three epochs, same batch, same optimiser, same seed, same augmentation:

| | Parameters | Test accuracy | |---|---|---| | TinyVGG | 0.13M | 68.30% | | ResNet18, ImageNet stem | 11.18M | 66.69% | | ResNet18, CIFAR stem | 11.17M | 78.79% |

Slightly fewer parameters, because a 3×3×3×64 kernel is smaller than 7×7×3×64.

Where the gain lands, and what it says

The per-class breakdown is the part I find convincing. The fix is concentrated in exactly the classes that need fine spatial detail: cat goes from 27.5% to 68.4%, dog from 59.9% to 73.8%. Classes with strong global colour and shape cues — automobile, ship, airplane — barely move, because 8×8 was already enough for them.

The confusions change character too. The broken model's worst pair was cat → frog, which is a mistake at the level of colour and texture. The fixed model's is cat → dog — the confusion CIFAR-10 models are supposed to have.

Three epochs is a small budget and 78.79% is not a strong absolute CIFAR-10 number. That is not what the exercise measures. It measures one change, with everything else pinned, which is the only way the +12.1 means anything.

Screens

Test accuracy for three models trained with identical settings: TinyVGG at 68.3% with 0.13M parameters, ResNet18 with the ImageNet stem at 66.7% with 11.18M parameters, and ResNet18 with a CIFAR stem at 78.8% with 11.17M parameters
Per-class accuracy for the three models across all ten CIFAR-10 classes, sorted ascending. The stem fix helps most on cat, dog and bird; cat rises from 27.5% to 68.4%