The Metadata Nobody Sees

Head of Product Design

Solo, with the logging team

A media catalogue is only as navigable as the metadata someone typed in by hand. This is about the hand.

The shape of the problem, without the domain

Every media company has metadata requirements that are specific to it, and specific enough that no off-the-shelf tool captures them.

Read the reasoning behind this decision

A cooking channel needs the ingredient list and the moment each technique is demonstrated. A conference organiser needs the speaker, the slide transitions, and the point where the Q&A starts. A sports broadcaster needs the goals. In every case there is a set of references inside the video that matter to that particular audience, and someone has to find them and write down where they are.

At Bibel TV the references are scripture. The catalogue is sermons, Bible studies and lectures, and the people watching want to know which passage is being discussed and when. Substitute your own domain and the problem is identical: timecoded references, extracted by a person watching in real time, because the thing that makes them valuable is exactly the thing that makes them hard to automate — they require understanding what was said, not just hearing it.

That is the general case. The rest of this is the specific one, and the specifics are where the design decisions live.

Two iterations, and what changed between them

Iteration one worked from the transcript. A transcript tells you reliably where a reference is *mentioned*. That is a different question from where the passage is *taken up*, and the gap between them is exactly the gap between a correct timestamp and a useful one.

Iteration two analyses the video itself. It locates a more meaningful start point, because it can attend to more than the words — the structural signals a transcript flattens. The comparison I am making is against my own earlier version, not against a vendor or a human baseline. That is the only comparison I have earned.

The system now attempts the whole capture, not just the references: structure, segments, graphics, cover frame. My estimate is that the complete system removes ten to fifteen minutes of work per video. It is in progress, not finished.

The decision I keep making

The team still watches the video. That is deliberate, and it is the part of the design I would defend hardest.

Read the reasoning behind this decision

What the system removes is the logging — the transcription, the timestamping, the screenshotting, the typing. What it leaves is the watching and the judgment: the guidelines and editorial rules that are harder to automate and more consequential when they are wrong. A person arrives at a new video with the metadata already extracted and generated, and spends their attention on the part that needed a person.

I have now made this same call three times without noticing I was repeating myself. The customer-care copilot at Bibel TV is arranged so a person reads the viewer's letter before the model ever sees it, and every send decision stays human. The AI onboarding model I designed at Ninox generates a database schema into the real data-model editor rather than a chat preview, so a person can see and correct what the model produced.

Automate the capture. Leave the judgment. Three systems, one principle, arrived at separately each time.

Want to hear more?

  • Curious why the fix for a search result titled "Who is Jesus?" was to give the clip its own name rather than a smarter search, or what a week logging videos by hand taught me about which parts of the job were never meant to be automated?
Get in touch