Runs 100% on your own machine

One long video in.
A batch of Shorts out.

Clipper Studio transcribes your podcast, interview, stream VOD or gameplay recording, asks a local LLM which moments are worth cutting, reframes each one to 9:16 while following whoever is talking, burns in karaoke subtitles, writes the metadata — and uploads it if you want.

MIT licensed Python 3.10+ No paid API Windows · Linux · macOS

Clipper Studio dashboard showing the New render panel with source video, URL and file upload options
How it works

Nine small agents, one at a time

A manager runs each stage in order with retries and per-project logging. Every stage writes its output to disk, so you can inspect, fix or re-run any part of it by hand.

01
Transcriptfaster-whisper, word-level timestamps
02
AnalyzerLLM scores highlight ranges
03
EditorCut + speaker-aware 9:16 crop
04
SubtitleKaraoke ASS, burned into the render
05
SEOTitle, description, hashtags
06
ThumbnailBest frame + title stamp
07
QCDuration, resolution, playable
08
ComplianceWord list + LLM policy scan
09
UploadYouTube, private by default
Features

Everything the pipeline gives you

Built for people who cut a lot of clips and got tired of doing the same nine steps by hand.

Five ways to get video in

A file from incoming/, one or more video or livestream URLs, a browser upload, a Google Drive folder, or any path on disk.

Speaker-aware reframing

YuNet face detection plus audio-visual speaker detection, so the vertical crop follows whoever is actually talking instead of the biggest face.

Karaoke subtitles

Word-by-word ASS subtitles burned into the render, with font, size, colour, outline and vertical margin all configurable.

Metadata that is ready to post

Title, description, hashtags and keywords per clip, in the language you pick — shaped by a persistent campaign brief if you have one.

Pause and stop that behave

Cooperative checkpoints between agents, clips and segments — a stop never leaves you with a half-written render.

Tunable without touching code

Clip length, crop tightness, loudness, silence trimming, subtitle style and the agent prompts themselves — all editable from the dashboard.

Multi-channel upload

Each YouTube channel keeps its own OAuth token. Bulk-select clips and send them out — only the ones that passed QC and compliance.

Talk to it in plain language

"Clips 10–30s, zoom out a bit, then run tes1.mp4" — the built-in assistant changes settings, applies briefs and starts renders.

The dashboard

A local web UI for the whole pipeline

One Flask app on 127.0.0.1:5000. Screenshots below use a small synthetic demo dataset, so the projects and clips shown are placeholders.

Review before anything goes out

Every clip shows its thumbnail, the score the LLM gave it, the generated title, duration, QC and compliance status, and whether it is already published. Select any combination and upload or delete them together.

bulk selecteditable SEOQC badges
Project detail page showing a grid of generated vertical clips with scores, titles and publish status

Place a watermark by dragging it

Upload a logo or add a text handle, then drag it onto a real frame from your latest render. The yellow zone marks where subtitles will sit, so you do not cover them. Watermarks can also be applied to an already-rendered project.

live previewmulti-watermarkre-apply
Watermark placement panel with a draggable preview frame and the subtitle safe zone highlighted

Tune the framing until it looks right

How tightly the 9:16 crop holds a face, whether it follows the active speaker, what happens when several people talk at once — plus loudness normalisation, silence trimming and punch-in motion.

YuNetAV speaker detect−14 LUFS
Framing settings with zoom slider, speaker tracking toggle and advanced detector options

More of the interface

Natural-language control over settings, briefs and renders
Stage-by-stage progress with a live log tail
Every render, searchable, pinned first
One OAuth token per YouTube channel
Send footage from anywhere on disk, without copying
Rewrite the agent prompts, with placeholder validation
Quick start

Running in about five minutes

Assuming ffmpeg and Ollama are already on your machine. Full details are in the README.

  1. Install the prerequisites

    Python 3.10+, ffmpeg and ffprobe on your PATH, and Ollama with a model pulled.

    ollama pull qwen2.5:7b
  2. Clone and install

    A virtualenv is recommended but not required.

    git clone https://github.com/dhimasbagus402/clipper-bot.git
    cd clipper-bot
    pip install -r requirements.txt -r requirements-dashboard.txt
  3. Create your config

    Both files are git-ignored — config.yaml holds your settings, .env your dashboard password.

    cp config.example.yaml config.yaml
    cp .env.example .env
  4. Start the dashboard

    Then open 127.0.0.1:5000 and drop in a video.

    python dashboard.py

What you need

Python3.10 or newer
ffmpegffmpeg + ffprobe on PATH
OllamaRunning locally, one model pulled
GPUOptional — CUDA speeds up transcription
DiskRenders are large; plan for several GB per hour of source

Prefer Docker?

The compose file mounts the project directory, so config, projects and tokens stay on the host. Ollama is expected to run on the host, not in the container.

docker compose up -d --build

No dashboard needed

The pipeline also runs headless.

python run.py incoming/podcast.mp4
Before you rely on it

What this does not do

Worth reading before you point it at a client campaign.

Copyright is not checked.

The compliance agent scans language, nothing else. It has no idea whether your music or footage is licensed, and YouTube's Content ID still applies to everything you upload.

Highlight scores are relative, not predictive.

The LLM ranks moments within one video. A 94 does not mean the clip will perform — it means the model liked it more than the 79.

Reframing is good, not perfect.

Speaker tracking works well for talking-head content. Fast-moving subjects and crowded frames still produce awkward crops.

YouTube uploads only.

The TikTok, Instagram and Facebook tiles in the interface are placeholders for planned work, not working integrations.

FAQ

Common questions

Do I need a GPU?

No. An NVIDIA GPU with CUDA makes transcription several times faster, but everything runs on CPU as well — it is simply slower on long videos.

Is my video uploaded to a cloud service?

No. Transcription uses faster-whisper and highlight selection uses Ollama, both running locally. The only network calls are the ones you ask for: downloading a source video, or uploading a finished clip to YouTube.

Does it check copyright or music licensing?

No. The compliance agent scans language only. It has no knowledge of whether your music or footage is licensed, and YouTube's Content ID still applies to anything you upload.

How many clips can I upload per day?

One upload costs roughly 1,600 units of YouTube's default 10,000/day quota, so about six uploads per day unless you request more from Google.

Can I expose the dashboard on my network?

Yes, but set DASH_PASSWORD first. The dashboard can browse your filesystem, start renders and upload to your account, so it must not be left open.

Can I change what the AI is told to do?

Yes. The system and user prompts for the Analyzer, SEO and Compliance agents are editable from Settings, with validation so you cannot save a prompt that breaks at render time.

Support the project

Clipper Studio is free and open source. If it saves you time, you can buy me a coffee.