Learning project · Sole Author · Apr 2026 — Apr 2026
Shelf Scanner
Counting objects that keep flickering
Study project, never deployed
What mattered
Counting happens on a band with a dwell, not a line, so a jittering box is not counted twice.
Context
An iOS app that counts stock on a shelf by pointing the camera at it. A YOLO model runs on-device through CoreML, detections are tracked frame to frame, and objects are counted as they pass through a band in the middle of the screen.
The recording above is a real device. Watch the count climb as the phone pans — that, and not the detection, is where the work went.
Detection is the easy half
Boxes on a still image are close to free. The problem is that a detector run on a moving camera does not produce a stable set of objects — it produces a slightly different set every frame. A can at an angle drops below threshold and vanishes for three frames. A hand passes over the shelf. The same can gets two overlapping boxes.
Counting on top of that is where it gets hard, because every one of those flickers is a potential double-count or a miss.
Count on a band, with a dwell
The obvious approach is a vertical line at x=0.5 and a count when a box's centroid crosses it. That breaks immediately on jitter: a centroid oscillating a pixel either side of the line counts the same can repeatedly.
So the counter uses a region, not a line:
private let countingRegionLeft: CGFloat = 0.45
private let countingRegionRight: CGFloat = 0.55
private let minDwellFrames: Int = 2
An object has to sit inside the band for two consecutive frames before it
counts, and a Set<Int> of already-counted track IDs makes each identity
countable exactly once. Leaving the band resets that track's dwell. The two blue
lines in the recording are the edges of this band.
One bug from that area is worth repeating because of how it hides: the overlay and the counting logic disagreed about where the line was, so the app counted at a different place from the one it drew. Everything looked correct. It has its own commit — align counting line UI with CountingLogic.
Recovering an object that dimmed instead of leaving
Counting once per track ID only works if the track ID survives. A can that is briefly occluded doesn't disappear from the model's output — it drops below the confidence threshold. Discard those weak detections and the track dies, a new one is minted when the can reappears, and the same can is counted twice.
That is what the two-stage association is for. Each frame:
- Stage 1 matches high-confidence detections against all
trackedandlosttracks. - Stage 2 matches the low-confidence detections against only the lost tracks stage 1 failed to match.
The cost is 1 - IoU(kalmanPredictedBox, detectionBox), optionally scaled by
(1 - confidence). Using the Kalman prediction rather than the last known box
matters while the camera is panning — by the time the next frame arrives the can
has moved, and IoU against a stale box is near zero.
The matcher is greedy, deliberately
An earlier Hungarian implementation was buggy on rectangular inputs — and the matrix is rectangular on essentially every frame, because the number of tracks and the number of detections rarely agree. I replaced it with a greedy matcher: for each detection, take the lowest-cost unused track under threshold.
Greedy is not globally optimal. It is O(n·m), allocation-free, and cannot fall over when tracks ≠ detections — which is the right trade at 30fps on a phone, where a dropped frame is more visible than a slightly suboptimal assignment.
(The function is still named hungarianAlgorithm while its body does greedy
matching. That is a misleading name and should be fixed.)
Thresholds, and a benchmark I would not quote
The detector's confidence threshold came down from 0.75 to 0.6, because 0.75 was missing cans that were partly occluded or at an angle, with soft-NMS at IoU 0.65 to collapse duplicate boxes. The direction is right for counting specifically: a missed detection is worse than a duplicate, because NMS and the tracker each remove duplicates, and nothing recovers a detection that never happened.
I also ran a threshold sweep, and I do not trust its timings. It ran through Ultralytics on the Mac's GPU rather than on-device CoreML, so its FPS numbers say nothing about the phone. Worse, its input-size results are impossible:
320x320 → 24.4 FPS 640x640 → 57.1 FPS
Smaller input, less than half the throughput. That is a warm-up artefact, not a
finding — the last configuration measured got a hot pipeline. The top five
conf/IoU rows land within 63–68 FPS of each other for the same reason, which is
noise. conf=0.6 is justified by what it detects, not by that ranking.
The model is trained on 400 images of a single class, so this is a working demonstration of the counting pipeline rather than a claim about generalising to other products or other lighting.
Screens
