Category: Behind the Build

  • We built our own video player

    The easy way to put a film on a web page is to paste in an embed code from somewhere else. We did not do that. Every film here plays inside a player we wrote, on our own page, with nothing redirecting anywhere.

    Underneath it is a plain HTML5 <video> element with the browser controls deliberately switched off. Everything you see — the play button, the clock, the scrub bar, the volume slider, fullscreen — is jQuery talking to the HTML5 media API.

    The interesting parts were the ones nobody notices. The timeupdate event fires several times a second, so the clock has to update without fighting you while you are dragging the scrub bar. The film’s duration is unknown until metadata loads, so the slider cannot be sized until then. Fullscreen has to apply to the wrapper rather than the video, or the controls disappear with it.

    It would have been faster to use somebody else’s player. We would also have learned nothing, and we would not be able to explain a single line of it.

  • How the live comment stream works

    Open any film in two browsers and type a comment in one. It appears in the other within about three seconds, and the viewer count updates at the same time. Here is what is actually happening.

    Each open page asks the server a question every three seconds: is there anything newer than comment number N, and who else is here? The browser sends the highest comment id it already holds, so the server replies with only what is new — usually nothing at all, which costs a few hundred bytes.

    Comments and the viewer count are served by one request, not two. They refresh on the same schedule and appear in the same panel, so splitting them would have doubled the traffic for no benefit.

    The count works because every browser generates a random token and sends it with each poll. The server records when it last heard from each token and forgets any that go quiet for thirty seconds. Closing a tab sends nothing, so silence is the only signal available.

    Polling is not the most elegant answer — WebSockets would push instead of asking. But polling works everywhere, needs no extra server process, and for a comment stream three seconds late is not late at all.

  • Where the catalogue comes from

    Twenty-one films did not get typed in by hand. They were built from the Internet Archive’s public API, and the process taught us something we did not expect.

    The Archive exposes a metadata endpoint for every item, listing each file it holds. That matters because the obvious assumption — that a film’s video file is named after the item — is frequently false. Guessing the URL would have produced a catalogue of broken links. Reading the file list instead gave us a verified address for every single film.

    The surprise came from how we chose them. Our first query sorted by download count, on the reasonable theory that popular films are good films. What came back was mostly exploitation cinema, with a reel of Holocaust documentation footage in the middle of it.

    Popularity turned out to be the wrong question. We threw the ranking away and chose every title deliberately instead. An archive is not a pile of files that happen to be free — the choosing is the work. That is the difference between storage and a collection.