Amir TabatabaeiWork with me
← All work

Learning project · Sole Author · Apr 2026 — Apr 2026

Shelf Scanner

Counting objects that keep flickering

Study project, never deployed

SwiftCoreMLYOLOByteTrackAVFoundation

What mattered

Counting happens on a band with a dwell, not a line, so a jittering box is not counted twice.

Context

An iOS app that counts stock on a shelf by pointing the camera at it. A YOLO model runs on-device through CoreML, detections are tracked frame to frame, and objects are counted as they pass through a band in the middle of the screen.

The recording above is a real device. Watch the count climb as the phone pans — that, and not the detection, is where the work went.

Detection is the easy half

Boxes on a still image are close to free. The problem is that a detector run on a moving camera does not produce a stable set of objects — it produces a slightly different set every frame. A can at an angle drops below threshold and vanishes for three frames. A hand passes over the shelf. The same can gets two overlapping boxes.

Counting on top of that is where it gets hard, because every one of those flickers is a potential double-count or a miss.

Count on a band, with a dwell

The obvious approach is a vertical line at x=0.5 and a count when a box's centroid crosses it. That breaks immediately on jitter: a centroid oscillating a pixel either side of the line counts the same can repeatedly.

So the counter uses a region, not a line:

private let countingRegionLeft:  CGFloat = 0.45
private let countingRegionRight: CGFloat = 0.55
private let minDwellFrames: Int = 2

An object has to sit inside the band for two consecutive frames before it counts, and a Set<Int> of already-counted track IDs makes each identity countable exactly once. Leaving the band resets that track's dwell. The two blue lines in the recording are the edges of this band.

One bug from that area is worth repeating because of how it hides: the overlay and the counting logic disagreed about where the line was, so the app counted at a different place from the one it drew. Everything looked correct. It has its own commit — align counting line UI with CountingLogic.

Recovering an object that dimmed instead of leaving

Counting once per track ID only works if the track ID survives. A can that is briefly occluded doesn't disappear from the model's output — it drops below the confidence threshold. Discard those weak detections and the track dies, a new one is minted when the can reappears, and the same can is counted twice.

That is what the two-stage association is for. Each frame:

  • Stage 1 matches high-confidence detections against all tracked and lost tracks.
  • Stage 2 matches the low-confidence detections against only the lost tracks stage 1 failed to match.

The cost is 1 - IoU(kalmanPredictedBox, detectionBox), optionally scaled by (1 - confidence). Using the Kalman prediction rather than the last known box matters while the camera is panning — by the time the next frame arrives the can has moved, and IoU against a stale box is near zero.

The matcher is greedy, deliberately

An earlier Hungarian implementation was buggy on rectangular inputs — and the matrix is rectangular on essentially every frame, because the number of tracks and the number of detections rarely agree. I replaced it with a greedy matcher: for each detection, take the lowest-cost unused track under threshold.

Greedy is not globally optimal. It is O(n·m), allocation-free, and cannot fall over when tracks ≠ detections — which is the right trade at 30fps on a phone, where a dropped frame is more visible than a slightly suboptimal assignment.

(The function is still named hungarianAlgorithm while its body does greedy matching. That is a misleading name and should be fixed.)

Thresholds, and a benchmark I would not quote

The detector's confidence threshold came down from 0.75 to 0.6, because 0.75 was missing cans that were partly occluded or at an angle, with soft-NMS at IoU 0.65 to collapse duplicate boxes. The direction is right for counting specifically: a missed detection is worse than a duplicate, because NMS and the tracker each remove duplicates, and nothing recovers a detection that never happened.

I also ran a threshold sweep, and I do not trust its timings. It ran through Ultralytics on the Mac's GPU rather than on-device CoreML, so its FPS numbers say nothing about the phone. Worse, its input-size results are impossible:

320x320 → 24.4 FPS      640x640 → 57.1 FPS

Smaller input, less than half the throughput. That is a warm-up artefact, not a finding — the last configuration measured got a hot pipeline. The top five conf/IoU rows land within 63–68 FPS of each other for the same reason, which is noise. conf=0.6 is justified by what it detects, not by that ranking.

The model is trained on 400 images of a single class, so this is a working demonstration of the counting pipeline rather than a claim about generalising to other products or other lighting.

Screens

A screen recording of the iOS app: green boxes with confidence labels track Coke cans on a shelf as the phone pans across it, two blue vertical lines mark the counting band in the middle of the screen, and the live count rises from three to six