Learning project · Sole Author · May 2025 — Aug 2026
CIFAR-10 Classifiers
An 86× larger network that lost
Study project, never deployed
What mattered
The ImageNet stem throws away a 32×32 image before the first residual block; the CIFAR stem keeps it.
System map
Context
Two classifiers trained from scratch on CIFAR-10: a hand-written TinyVGG and
torchvision's ResNet18. Three epochs each, batch 32, Adam at 1e-3, the same
flip-and-crop augmentation.
Both checkpoints were still in the repo a year later, so I could measure them rather than trust the training logs. The result is the reason this is worth a page:
| | Parameters | Test accuracy | |---|---|---| | TinyVGG | 0.13M | 68.30% | | ResNet18 | 11.18M | 66.69% |
Eighty-six times the parameters, and it loses. Same budget, same data, same augmentation — so it is not undertraining.
The stem throws the image away
torchvision's ResNet18 is built for 224×224 ImageNet input. Its stem is a 7×7
stride-2 convolution followed by a 3×3 stride-2 maxpool — a 4× linear reduction
before any residual block runs. On a 224×224 image that is sensible. On a 32×32
image it is most of the picture.
Rather than argue it, I measured the tensor:
ResNet18 ImageNet stem: 32x32 -> (8, 8) entering layer1
ResNet18 CIFAR stem: 32x32 -> (32, 32) entering layer1
Eight by eight. Then layer2 through layer4 downsample three more times, to
1×1. Eleven million parameters are operating on an image that no longer exists,
which is why a 0.13M-parameter network that kept its resolution beat it.
Two lines, +12.1 points
The standard CIFAR adaptation replaces the stem and drops the maxpool:
m.conv1 = nn.Conv2d(3, 64, kernel_size=3, stride=1, padding=1, bias=False)
m.maxpool = nn.Identity()
I retrained with everything else held fixed — same three epochs, same batch, same optimiser, same seed, same augmentation:
| | Parameters | Test accuracy | |---|---|---| | TinyVGG | 0.13M | 68.30% | | ResNet18, ImageNet stem | 11.18M | 66.69% | | ResNet18, CIFAR stem | 11.17M | 78.79% |
Slightly fewer parameters, because a 3×3×3×64 kernel is smaller than 7×7×3×64.
Where the gain lands, and what it says
The per-class breakdown is the part I find convincing. The fix is concentrated
in exactly the classes that need fine spatial detail: cat goes from 27.5% to
68.4%, dog from 59.9% to 73.8%. Classes with strong global colour and shape
cues — automobile, ship, airplane — barely move, because 8×8 was already
enough for them.
The confusions change character too. The broken model's worst pair was
cat → frog, which is a mistake at the level of colour and texture. The fixed
model's is cat → dog — the confusion CIFAR-10 models are supposed to have.
Three epochs is a small budget and 78.79% is not a strong absolute CIFAR-10 number. That is not what the exercise measures. It measures one change, with everything else pinned, which is the only way the +12.1 means anything.
Screens

