← Back to portfolio
CASE STUDY · BIBEL TV · 2026 – PRESENT

The Metadata Nobody Sees

Head of Product Design

Solo, with the logging team

A media catalogue is only as navigable as the metadata someone typed in by hand. This is about the hand.

What the job actually was

In broadcast and post-production this work has a name: logging. A logger reviews footage and records timecoded descriptions of what is in it. At Bibel TV, for every newly ingested video, a member of the logging team had to:

- watch it, completely — not skim, not scrub - capture the general timeline and structure, identifying the distinct segments - timestamp each of them - screenshot every graphic: full-screen cards, lower thirds, logos - extract and document the scripture references - capture a frame to use as the cover image

That is the core of it rather than the whole of it, and it is one pass, per asset, forever. None of it is visible to a viewer. All of it determines whether the catalogue can be searched, browsed or recommended at all.

Nothing on that list is unique to a faith broadcaster. Any media organisation with a catalogue worth navigating employs loggers to do some version of it.

I know the shape of that work because I did it. I spent a week on the logging team as part of my onboarding, doing the capture by hand on real videos. That week is the reason I have a number to quote and the reason I trust it: by my own count the scripture extraction alone took three to five minutes per video, sometimes more. I would call that a conservative estimate rather than a measured study, and I would rather say where it came from than dress it up as data.

The decision I keep making

The team still watches the video. That is deliberate, and it is the part of the design I would defend hardest.

What the system removes is the logging — the transcription, the timestamping, the screenshotting, the typing. What it leaves is the watching and the judgment: the guidelines and editorial rules that are harder to automate and more consequential when they are wrong. A person arrives at a new video with the metadata already extracted and generated, and spends their attention on the part that needed a person.

I have now made this same call three times without noticing I was repeating myself. The customer-care copilot at Bibel TV is arranged so a person reads the viewer's letter before the model ever sees it, and every send decision stays human. The AI onboarding model I designed at Ninox generates a database schema into the real data-model editor rather than a chat preview, so a person can see and correct what the model produced.

Automate the capture. Leave the judgment. Three systems, one principle, arrived at separately each time.

Read the full case study

The shape of the problem, without the domain

Every media company has metadata requirements that are specific to it, and specific enough that no off-the-shelf tool captures them.

A cooking channel needs the ingredient list and the moment each technique is demonstrated. A conference organiser needs the speaker, the slide transitions, and the point where the Q&A starts. A sports broadcaster needs the goals. In every case there is a set of references inside the video that matter to that particular audience, and someone has to find them and write down where they are.

At Bibel TV the references are scripture. The catalogue is sermons, Bible studies and lectures, and the people watching want to know which passage is being discussed and when. Substitute your own domain and the problem is identical: timecoded references, extracted by a person watching in real time, because the thing that makes them valuable is exactly the thing that makes them hard to automate — they require understanding what was said, not just hearing it.

That is the general case. The rest of this is the specific one, and the specifics are where the design decisions live.

What the references are actually for

The extracted references are not decoration on a video record. They run in two directions, and the second one is the product.

Video to scripture. Watching a sermon, you can open the passage being discussed in the app's Bible feature, at the point it is discussed.

Scripture to video. Reading a passage in the Bible feature, you can see every video in the catalogue that discusses it.

That second direction turns a verse into an index across the entire library. It is also what makes the timecode a jump target rather than a note — someone taps a result and lands inside a video. A technically correct timestamp that drops you mid-sentence, or after the point has been made, is a worse product than no timestamp at all.

This is the distinction the first version of the system got wrong, and the reason there is a second one.

Two iterations, and what changed between them

Iteration one worked from the transcript. A transcript tells you reliably where a reference is *mentioned*. That is a different question from where the passage is *taken up*, and the gap between them is exactly the gap between a correct timestamp and a useful one.

Iteration two analyses the video itself. It locates a more meaningful start point, because it can attend to more than the words — the structural signals a transcript flattens. The comparison I am making is against my own earlier version, not against a vendor or a human baseline. That is the only comparison I have earned.

The system now attempts the whole capture, not just the references: structure, segments, graphics, cover frame. My estimate is that the complete system removes ten to fifteen minutes of work per video. It is in progress, not finished.

The addition I am introducing: virtual clips

The existing system records a start time for each reference. I am adding the end time — the span over which the video actually discusses that passage.

A reference stops being a point and becomes a segment. A segment can be addressed. And an addressable segment can carry its own metadata, which is where this stops being a pipeline improvement and becomes a content-model change.

Here is the concrete case. A sermon about Jesus also discusses the creation account. Today, someone reading the creation account in the Bible feature gets a result titled "Who is Jesus?" The passage genuinely is discussed in there. The title conceals it. An inherited title describes the parent video, not the reason this result surfaced — so a relevant hit reads as an irrelevant one, and a feature that is working correctly looks broken.

Give the segment its own title — *"What Jesus has to do with creation"* — and it argues for its own relevance. That is the intent: better discovery, and visible relevance inside the Bible feature. It is not shipped, and I have not measured it. When it ships, that is the number worth having.

What I took from it

The week in the department is the part I would repeat anywhere. Not because it produced a number, though it did, but because it produced the right instinct about which minutes were worth removing. Sitting with the workflow told me that the tedium was in the logging and the value was in the watching — and that is not a conclusion I would have reached from a requirements document or a stakeholder interview.

The other thing I would carry forward is smaller and more specific. A timestamp is an interface. It looked like a data field for as long as I thought of it as a record of where something occurs. It became a design problem the moment I understood that a person taps it and arrives somewhere, and that arriving in the wrong place is a worse experience than not being offered the jump.

Most of the metadata nobody sees is like that. It reads as plumbing until you find the surface where a person meets it.

Curious why the fix for a search result titled "Who is Jesus?" was to give the clip its own name rather than a smarter search, or what a week logging videos by hand taught me about which parts of the job were never meant to be automated? Get in touch.