DarkStyleee 9575612f06 transcribe: select the audio track, and refuse to upload a silent one (#134)
* feat(transcribe): pick the audio track and refuse to upload silence

extract_audio ran without -map, so ffmpeg applied its default stream selection
and took a single audio track — the first one. A multi-track recording is the
normal case for a screen capture: OBS writes the application on track 0 and the
microphone on track 1. Transcribing such a file uploaded the application audio
and dropped the narration without a word about it.

--audio-track selects the stream, zero-based, and defaults to 0, so existing
single-track runs are unchanged. A file with more than one track says so in
verbose output, since the default is right for some of them and wrong for others.

The extracted wav is also checked for level before it is sent. A peak under
-60 dBFS means the track is silent, which in practice means the wrong track was
picked, and Scribe charges the same for silence as for speech. The error names
the track count and points at the flag.

* fix(transcribe): key the cache by track, scan the peak in chunks, name the tracks

Review on #134 caught four things.

The transcript cache is keyed by video stem alone, and transcribe_one returns on
a cache hit before it looks at audio_track. Rerunning with --audio-track 1 after
a wrong-track run therefore handed back the very transcript the flag was meant to
replace. The track goes into the file name now, and track 0 keeps the old name so
existing transcripts stay valid.

peak_dbfs read the whole take with readframes(getnframes()) and copied it into an
array, so a two-hour 16 kHz mono file cost 230 MB twice over, with batch mode
running several at once. It scans in 64k-frame chunks instead; measured on a real
capture the peak is identical to the whole-file version.

The "try --audio-track" hint flipped between 0 and 1, so on a file with three
tracks it could point at another silent one. It lists the tracks that exist:
"The file has 3 audio tracks; try --audio-track 0 or 2."

The flag's help text said ffmpeg would otherwise take track 0. It applies its
default stream selection, which picks the track with the most channels.

* fix(transcribe): share one transcript path between single and batch mode

Review on #134 again: the previous commit put the track into the cache key in
transcribe.py and left transcribe_batch.py testing for {stem}.json. Batch mode
with --audio-track 1 therefore counted a file with a track-0 transcript as
cached and skipped it, which defeats the flag, and never recognised the
{stem}.track1.json it had just written, so it re-uploaded and re-billed that
file on every run.

Both now call transcript_path(), so the two cannot drift apart again.

Verified on a directory holding one video and a track-0 transcript: the default
run reports "1 cached, 0 to transcribe", the same run with --audio-track 1
reports "0 cached, 1 to transcribe" and goes on to the silence guard.
2026-08-30 02:45:51 -07:00
v1
2026-04-11 18:34:21 -07:00
2026-04-15 15:47:06 -07:00
v1
2026-04-11 18:34:21 -07:00
v1
2026-04-11 18:34:21 -07:00
2026-05-10 11:56:51 -07:00
2026-04-15 15:47:06 -07:00
2026-05-10 11:56:51 -07:00
2026-06-29 11:03:55 +08:00

video-use

video-use

Introducing video-use — edit videos with Claude Code. 100% open source.

Drop raw footage in a folder, chat with Claude Code, get final.mp4 back. Works for any content — talking heads, montages, tutorials, travel, interviews — without presets or menus.

Try video-use in Browser Use Cloud.

What it does

  • Cuts out filler words (umm, uh, false starts) and dead space between takes
  • Auto color grades every segment (warm cinematic, neutral punch, or any custom ffmpeg chain)
  • 30ms audio fades at every cut so you never hear a pop
  • Burns subtitles in your style — 2-word UPPERCASE chunks by default, fully customizable
  • Generates animation overlays via HyperFrames, Remotion, Manim, or PIL — spawned in parallel sub-agents, one per animation
  • Self-evaluates the rendered output at every cut boundary before showing you anything
  • Persists session memory in project.md so next week's session picks up where you left off

Setup prompt

Paste into Claude Code, Codex, Hermes, Openclaw, or any agent with shell access:

Set up https://github.com/browser-use/video-use for me.

Read install.md first to install this repo, wire up ffmpeg, register the skill with whichever agent you're running under, and set up the ElevenLabs API key — ask me to paste it when you need it. Then read SKILL.md for daily usage, and always read helpers/ because that's where the editing scripts live. After install, don't transcribe anything on your own — just tell me it's ready and wait for me to drop footage into a folder.

The agent handles the clone, dependencies, skill registration, and prompts you once for your ElevenLabs API key (grab one at elevenlabs.io/app/settings/api-keys).

Then point your agent at a folder of raw takes:

cd /path/to/your/videos
claude    # or codex, hermes, etc.

For always-on editing from your own VPS or Telegram, run the agent through Browser Use Box. Watch the 15-second demo.

And in the session:

edit these into a launch video

It inventories the sources, proposes a strategy, waits for your OK, then produces edit/final.mp4 next to your sources. All outputs live in <videos_dir>/edit/ — the skill directory stays clean.

Manual install

If you'd rather do it by hand:

# 1. Clone and symlink into your agent's skills directory
git clone https://github.com/browser-use/video-use ~/Developer/video-use
ln -sfn ~/Developer/video-use ~/.claude/skills/video-use        # Claude Code
# ln -sfn ~/Developer/video-use ~/.codex/skills/video-use       # Codex

# 2. Install deps
cd ~/Developer/video-use
uv sync                         # or: pip install -e .
brew install ffmpeg             # required
brew install yt-dlp             # optional, for downloading online sources

# 3. Add your ElevenLabs API key
cp .env.example .env
$EDITOR .env                    # ELEVENLABS_API_KEY=...

How it works

The LLM never watches the video. It reads it — through two layers that together give it everything it needs to cut with word-boundary precision.

timeline_view composite — filmstrip + speaker track + waveform + word labels + silence-gap cut candidates

Layer 1 — Audio transcript (always loaded). One ElevenLabs Scribe call per source gives word-level timestamps, speaker diarization, and audio events ((laughter), (applause), (sigh)). All takes pack into a single ~12KB takes_packed.md — the LLM's primary reading view.

## C0103  (duration: 43.0s, 8 phrases)
  [002.52-005.36] S0 Ninety percent of what a web agent does is completely wasted.
  [006.08-006.74] S0 We fixed this.

Layer 2 — Visual composite (on demand). timeline_view produces a filmstrip + waveform + word labels PNG for any time range. Called only at decision points — ambiguous pauses, retake comparisons, cut-point sanity checks.

Naive approach: 30,000 frames × 1,500 tokens = 45M tokens of noise. Video Use: 12KB text + a handful of PNGs.

Same idea as browser-use giving an LLM a structured DOM instead of a screenshot — but for video.

Pipeline

Transcribe ──> Pack ──> LLM Reasons ──> EDL ──> Render ──> Self-Eval
                                                              │
                                                              └─ issue? fix + re-render (max 3)

The self-eval loop runs timeline_view on the rendered output at every cut boundary — catches visual jumps, audio pops, hidden subtitles. You see the preview only after it passes.

Design principles

  1. Text + on-demand visuals. No frame-dumping. The transcript is the surface.
  2. Audio is primary, visuals follow. Cuts come from speech boundaries and silence gaps.
  3. Ask → confirm → execute → self-eval → persist. Never touch the cut without strategy approval.
  4. Zero assumptions about content type. Look, ask, then edit.
  5. 12 hard rules, artistic freedom elsewhere. Production-correctness is non-negotiable. Taste isn't.

See SKILL.md for the full production rules and editing craft.

S
Description
Edit any video by conversation. Transcribe, cut, color grade, generate overlay animations, burn subtitles — for talking heads, montages, tutorials, travel,…
Readme MIT 35 MiB
Languages
Python 79.7%
HTML 19.4%
Shell 0.9%