Where the catalogue comes from

Twenty-one films did not get typed in by hand. They were built from the Internet Archive’s public API, and the process taught us something we did not expect.

The Archive exposes a metadata endpoint for every item, listing each file it holds. That matters because the obvious assumption — that a film’s video file is named after the item — is frequently false. Guessing the URL would have produced a catalogue of broken links. Reading the file list instead gave us a verified address for every single film.

The surprise came from how we chose them. Our first query sorted by download count, on the reasonable theory that popular films are good films. What came back was mostly exploitation cinema, with a reel of Holocaust documentation footage in the middle of it.

Popularity turned out to be the wrong question. We threw the ranking away and chose every title deliberately instead. An archive is not a pile of files that happen to be free — the choosing is the work. That is the difference between storage and a collection.

Comments

Leave a Reply