Consulting Work Blog Contact
← Back to blog

The Atlas Hidden in Our Photos: What 67,608 Family Pictures Knew About Where We'd Been

The Atlas Hidden in Our Photos: What 67,608 Family Pictures Knew About Where We'd Been

I had 67,608 photos. Twenty years of them, the oldest dating to 2006, scattered across 73 different cameras and phones the family has owned since. They were all backed up, all “safe,” and completely unreachable. Nobody scrolls back through twenty years of pictures to find the afternoon in a particular town, because the act of finding is harder than the act of remembering. The archive wasn’t worthless. It was a second brain with no front door. So I built one, and pointed AI at the whole pile to see what was actually in there. The thing that came out the other end wasn’t a better search box. It was a set of maps I didn’t know I was carrying.

TL;DR

  • 67,608 photos across twenty years and 73 cameras, all given a machine-readable “sense” of what’s in them.
  • I taught a generic model my personal taste from about 1,000 of my own favourites — no giant training run, just a small head on top.
  • 53,242 photos already knew where they were taken. For the 14,363 that didn’t, a model guessed from the pixels alone — impressively, and unreliably.
  • The boring method (fill in a missing location from the photos taken minutes either side) beat the flashy one every time.
  • The payoff wasn’t a search box. It was 86 trips, two decades of them, drawn as keepsake maps from the trail the photos left behind.

Every family has this archive, and so does every company

This isn’t really a photos story. Every organisation I’ve ever seen is sitting on the same thing at a larger scale: two decades of email nobody can search by meaning, a shared drive with 40,000 documents and no map of them, years of support tickets that contain the answer to next week’s question if anyone could find it. The data is all there, all backed up, and functionally dead, because the cost of going through it by hand is higher than the value of any single thing you’d find. My camera roll was just the version of that archive that happened to be mine, and small enough that one person could try to wake it up in a weekend.

The goal was modest. I wanted twenty years of pictures to become something I could actually ask questions of. Where have we actually been. Which of these are any good. Show me the good ones from the coast, the year my daughter — well, the kind of question you can only answer if the pile knows what it contains. None of that is reachable in a folder of 67,608 anonymous files. All of it is sitting inside the images themselves, waiting for something to read them.

Step one: give every photo a sense of what it is

The first move is the one that makes everything else possible, and it’s quietly become cheap. You run every image through a vision model that turns the picture into a string of numbers — an “embedding” — that captures what the photo is about. Two pictures of a beach end up with similar numbers even if they were taken eleven years and two continents apart. An open, off-the-shelf model does this; I ran all 67,608 photos through it on my own machine, no cloud and no per-image fee.

Once every photo is a point in that space, “search” stops meaning filenames and starts meaning meaning. You can ask for “snowy mountain village” in plain words and get the right pictures back without anyone ever having tagged a single one. For a company this is the whole point: the embedding is the difference between an archive you grep and an archive you can actually talk to. The boring infrastructure step — turn everything into numbers once — is what makes the interesting questions possible later.

Step two: teach it my taste, not “good taste”

Here’s where it got personal, and where the most transferable lesson hides. There’s an off-the-shelf model that scores how “aesthetically pleasing” a photo is — it’ll happily tell you a sunset is prettier than a parking lot. Useful, and not what I wanted. Generic prettiness isn’t mine. My favourite photos are often technically unremarkable and personally irreplaceable.

So instead of accepting the generic score, I taught the system my taste. I took the roughly 1,000 photos I’d marked as favourites over the years as examples of “yes,” a pile of throwaway shots as examples of “no,” and trained a tiny model on top of the frozen embeddings. Not a giant training run — the heavy model stays exactly as it is — just a small head that learns the one thing it doesn’t know: what I like.

It worked better than I expected. On a scale of zero to one, the photos I’d hand-picked as favourites scored 0.84 on average. The library as a whole averaged 0.20. The model had genuinely learned the shape of my preference from a thousand examples, and could now apply it to all 67,608: 7,737 photos cleared a 0.8, and 5,380 cleared 0.9 — a hand-picked best-of I never had the patience to assemble myself. The lesson for anyone with a generic AI model and a specific need: you very rarely need to train your own giant. You need a frozen general model and a few hundred of your own labelled examples to teach it the one thing that’s yours.

Step three: the location nobody recorded

Then came the part that turned a search box into an atlas. Of the 67,608 photos, 53,242 already carried GPS coordinates baked in by the phone that took them. That alone is a map of twenty years. But 14,363 had no location at all — older cameras, stripped metadata, screenshots of moments.

So I tried to put them back on the map, and this is the part where the flashy thing and the boring thing went head to head. For 10,867 of them I used a model that guesses where on Earth a photo was taken from the pixels alone — no metadata, just the look of the light and the landscape. It’s genuinely astonishing that this works at all. It’s also rarely sure of itself, and honest about it: its average confidence was 0.04. It will tell you a beach is “probably Mediterranean” and be wrong often enough that you can’t trust any single guess.

The other 3,496 I filled in the dull way: if a photo has no location but the pictures taken three minutes before and after it do, it was almost certainly in the same place. The average confidence of this method was 0.74. It is unglamorous, it only works when neighbours exist, and it was right almost every time. So the flashy pixel-model went on the maps as a faint hint; the boring interpolation went on as fact. That gap — impressive but unsure versus dull but reliable — is the single most useful thing I relearned, and it generalises to every “look what the AI can do” demo you’ll ever be shown.

Twenty years, eighty-six trips, drawn from the trail

Once every photo had a location — recorded or recovered — the trail clustered itself into 86 distinct trips, twenty years of them. And a trip with coordinates and dates isn’t a list anymore. It’s a route. So I drew each one on a satellite map, the way you’d actually want to remember it.

A road trip across the American Southwest, reconstructed from the GPS in our photos: New York, Houston, San Antonio, Albuquerque, the Grand Canyon, Los Angeles
The American Southwest, 2017 — a route rebuilt from the GPS hidden in our photos.

This one started as a flight marker — “in from Moscow” into New York, then a hop down to Houston — and became three thousand kilometres of the American Southwest, the whole arc from the Gulf coast out to Los Angeles, traced out of the order the photos were taken. I didn’t keep an itinerary. The camera did.

Germany’s Romantic Road through Bavaria, drawn from our photos: Würzburg, Rothenburg, Dinkelsbühl, Nördlingen, the castles at Neuschwanstein, down to Füssen
Bavaria's Romantic Road — the same trip, replayed from our geotagged camera roll.

The Romantic Road, a loop out of home and back, every town we stopped in surfacing as a labelled dot because we happened to take a picture there. No planning document survives from that week. This one was reconstructed entirely from where the shutter clicked.

Western Turkey from the camera roll: Istanbul across to Bursa, down the Aegean coast through Izmir and Manisa
Western Turkey, Istanbul to the Aegean — another route recovered from photo metadata.

None of these maps existed a month ago, and all of the information in them existed all along. That’s the part I keep turning over. I didn’t add anything. The trips were already encoded in twenty years of photos; they were just unreadable until every picture knew where and when it was. The maps are the archive finally saying out loud what it had always known.

The before and after, in plain terms

BeforeAfter
Photos67,608 unsearchable filessearchable by meaning, in words
Qualityevery shot equal in the pilemy own taste, 5,380 top picks surfaced
Location53,242 mapped, 14,363 lostevery photo placed, recorded or recovered
Twenty yearsan undated heap86 trips, drawn as maps

What I’d do differently

I spent too long early on trusting the impressive thing. The pixel-based location model is a genuine marvel, and I wanted it to be the hero — to wave a wand over the 14,363 unlocated photos and hand them all back. I had to be talked out of that by its own confidence scores, which were quietly screaming “don’t trust me on any single one.” The honest read was there from the start: lean on the boring, reliable method wherever it applies, and demote the flashy one to a hint. I’d believe the numbers sooner next time instead of believing the demo.

The bigger takeaway, for anyone sitting on a dead archive and assuming it’s a lost cause, is that the economics flipped and most people haven’t noticed. Reading twenty years of unstructured material used to mean paying a human to go through it. Now the heavy lifting — understanding what each item is, scoring it against your own judgement, even recovering metadata that was never recorded — runs cheaply, often on your own machine. The archive was never worthless. It was just waiting for the cost of asking it a question to fall to nearly nothing, and that already happened.

Where would you start, knowing the one archive everyone wrote off can now be woken up in a weekend?

Stack: Python · EXIF · SQLite

Need something like this for your own business? See how I can help →