The engine watches your whole stream — the laughs, the reactions, the moment the room turns — and hands back finished vertical shorts, captioned and ready for TikTok, Reels and Shorts. One ten-hour stream gave 25 of them — in 91 minutes.
Gold marks every moment that cleared the bar — and they all become clips. You also get the score for the ones that didn’t.
A deliberately unflattering source — a 720p 30 fps recording, the kind most people actually have. Every clip still leaves at 1080 × 1920 with the captions checked pixel by pixel against the screen.
This is the workspace doing a real job, replayed in forty seconds: the stages are the engine’s own, in the proportions they took on the ten-hour stream; the clips at the end are the ones it delivered.
Everything heavy runs on our servers. Close the tab if you like — the work carries on without you, and nothing touches your machine.
MEASURED END TO END ON A 10-HOUR STREAM: 91 MINUTES, DOOR TO DOOR · NOTHING RUNS ON YOUR MACHINE
Every number below we measured on real footage and can prove. The methods behind them are ours.
Laughter, sudden energy, the moment several people talk at once — the engine hears the room turn and does not need the transcript to spell it out. Flat routine chatter is checked on every job and must score zero.
Separates speakers and gives each their own caption colour. When two voices overlap and it cannot be certain, it leaves the line uncoloured rather than colour it wrongly.
Switching language mid-sentence is expected, not an error case. The engine adapts to what is actually being talked about before it writes a word down.
Slurs are masked in text and bleeped in audio. Beyond that, moments that read as racist are never cut into a clip at all — because a bleep does not save a clip that is about the joke.
Word-by-word karaoke, several speakers in separate screen zones. We then sample the finished file and check the pixels — a caption that failed to render is caught before you ever see it.
Not a template. A quick beat stays short; a twelve-turn exchange gets room. Sentences always finish.
Monthly hours reset each month — extra hours you buy never expire. A failed job is never charged.
Several independent signals from the audio and the picture, weighted and scored across the whole recording. We publish the result rather than the recipe: every job comes with the score each moment got and the reason the ones we skipped lost. You can check the judgement on your own footage — that is what the two free hours are for.
Because a transcriber that guesses the wrong context does not degrade gracefully — it produces confident nonsense. The engine works out what is actually being talked about before it commits, which is why Swedish streams that slide into English come out readable instead of mangled.
The upload is stored for processing and deleted automatically within 7 days of delivery. We do not train models on your content. Voice profiles exist only if you create them and can be deleted at any time.
Yes. Pick who should be in focus, set preferred length, and extend the list of words that disqualify a moment entirely. Every choice the engine made is visible to you afterwards.
Short videos are back in a few minutes. Long streams run several times faster than real time — a ten-hour stream came back in 91 minutes. All of it runs on our servers — start the job and close the tab.
Upload an hour and read the detector output yourself. That is the whole pitch.
Clip engine v5 · measured, not guessed