Skip to content
jdmonacoPublic

About

Video frame capture and AI transcription pipeline for markdown notes

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

42 Commits

Folders and files

Repository files navigation

vidflow

Unified video capture and transcription CLI. Consolidates YouTube/local video frame extraction (formerly ytcapture) and AI vision transcription (formerly vidscribe) into a single installable package. Transcription runs on the local inference stack by default (ampere-gateway primary slot via ~/tools/aikit), with claude-* models as an Anthropic-API escape hatch.

Install

cd ~/tools/vidflow
uv sync

Or install as a tool:

uv tool install ~/tools/vidflow

This provides the vidflow command. The former standalone ytcapture, vidcapture, and vidscribe commands were removed in 0.6.0; use vidflow youtube, vidflow local, and vidflow transcribe.

Usage

YouTube capture

# Capture frames only
vidflow youtube https://youtube.com/watch?v=VIDEO_ID

# Bare video IDs also work
vidflow youtube dQw4w9WgXcQ

# No arguments: YouTube URLs are read from the clipboard (macOS),
# listed, and confirmed before capture (-y skips the prompt)
vidflow youtube

# Capture + full visual transcription in one step
vidflow youtube URL --transcribe

# Capture + text-only caption polish (cheaper: no frames sent)
vidflow youtube URL --polish

# Multiple videos (always independent — one note per video)
vidflow youtube URL1 URL2 --transcribe

# Playlist URLs expand to their videos (one note per video);
# watch URLs carrying a &list= param capture just that video
vidflow youtube "https://www.youtube.com/playlist?list=PLAYLIST_ID"

# Plan a run: which videos would be captured or skipped (already in the
# output directory). Playlists are expanded (one request each); no
# video is fetched and no model is called
vidflow youtube URL1 URL2 --polish --dry-run

Local video capture

# Capture frames from local file
vidflow local recording.mp4

# Capture + transcribe
vidflow local recording.mp4 --transcribe

# Multiple files merged
vidflow local part1.mp4 part2.mp4 --transcribe --merge

Caption text for a local capture comes from a sidecar WebVTT file when one sits next to the video, otherwise from an embedded text subtitle track. Sidecars match <stem>.vtt, <stem>.<lang>.vtt, or <stem>-<lang>.vtt (the last is how Teams/Stream name a downloaded transcript, e.g. talk-en-US.vtt), preferring the bare name, then English. A Teams meeting recording (<Meeting>-20261008_145749UTC-Meeting Recording.mp4) also matches a transcript named after the meeting alone (<Meeting>.vtt); since a recurring meeting reuses that name every time, such a match is skipped when its cues run more than a minute past the end of the video. --vtt FILE names the transcript explicitly for a single input, overriding discovery and embedded tracks. WebVTT voice tags (<v Speaker Name>, written by Teams meeting transcripts) become speaker turns: each turn is its own **Speaker Name**: text paragraph, and polish/transcribe keep those labels. --subtitle-track N selects an embedded track and skips the sidecar; --no-subtitles ignores all of them; --list-subtitles and --dry-run show which source would be used.

Transcribe existing captures

# Transcribe a single capture markdown
vidflow transcribe capture.md

# Several captures: one transcript each
vidflow transcribe talk1.md talk2.md

# Merge multiple captures into one transcript
vidflow transcribe part1.md part2.md --merge -o combined.md

# Estimate token usage before processing
vidflow transcribe capture.md --estimate-only

# Dry run
vidflow transcribe capture.md --dry-run

The transcript is written as a new note named from the generated title, and the input capture note moves into the sibling transcripts/ folder as the raw record (next to the raw-transcript JSON from capture), so each folder keeps one note per video. Pass --keep-capture to leave the capture note in place.

Polish existing captures (text-only)

polish is the lightweight alternative to transcribe: it sends only the collated caption text (YouTube auto-captions, or a local video's sidecar .vtt or embedded subtitles) to the configured model for cleanup — speech-to-text errors, filler words, punctuation, paragraphing — without sending any frame images. Sections without caption text pass through unchanged; frames-only captures are rejected (use transcribe).

Polish only improves its inputs: each file is polished on its own and in place — the section text is replaced while the file's frontmatter, title, and preamble (video embed, description) are preserved verbatim. Polish never merges inputs, retitles, or generates frontmatter. The raw captions remain recoverable in transcripts/raw-transcript-<id>.json. -o writes the polished copy elsewhere instead — a file for a single input, or a directory that receives each input under its own filename — and the raw capture note then moves to transcripts/ unless --keep-capture is given.

# Polish a capture markdown in place
vidflow polish capture.md

# Polish several captures, each in place
vidflow polish talk1.md talk2.md

# Write the polished copy elsewhere instead
vidflow polish capture.md -o polished.md
vidflow polish talk1.md talk2.md -o polished/

# Estimate token usage before processing
vidflow polish capture.md --estimate-only

Shell completion

vidflow completion bash --install   # Symlink into ~/.local/share/bash-completion/completions/
vidflow completion bash --path      # Show the installation path

The symlink targets the installed package, so re-run --install after reinstalling vidflow (e.g., switching between editable and regular installs).

Common options

# JSON output (stdout, for piping)
vidflow youtube URL --json

# Custom model
vidflow youtube URL --transcribe -m claude-opus-5   # Anthropic escape hatch

# Background context for transcription
vidflow transcribe capture.md -c agenda.md -c speakers.md

# Override title
vidflow transcribe capture.md -t "Workshop Day 1"

Environment

Variable Required Description
ANTHROPIC_API_KEY claude-* models only Anthropic API key for the escape-hatch lane
AMPERE_GATEWAY_URL Optional Gateway origin (default: http://ampere.lan:8080)
EXA_API_KEY No Enables citation search during transcription

Architecture

vidflow youtube URL --transcribe
  |
  +- vidflow.capture.core.process_video()      -> markdown with YouTube transcript
  |
  +- vidflow.youtube.transcribe_youtube()
       +- parse_vidcapture_markdown()           -> preserves existing transcript per section
       +- VidscribeProcessor.process_all()      -> AI vision transcription (local or Anthropic)

vidflow local file.mp4 --transcribe
  |
  +- vidflow.capture.core.process_local_video() -> markdown (sidecar/embedded captions, or empty sections)
  |
  +- vidflow.transcribe.transcribe_markdown()
       +- VidscribeProcessor.process_all()      -> standard skeleton transcription

When transcribing YouTube captures, existing auto-caption text is passed to the model via <existing-transcript> tags, instructing it to enhance and correct using visual frame context rather than transcribing from scratch.

vidflow polish capture.md
  |
  +- vidflow.transcribe.polish_markdown()
       +- parse_vidcapture_markdown()           -> collects caption text per section
       +- VidscribeProcessor(text_only=True)    -> text-only cleanup, no frames sent

Polish reuses the same processor, batching, retry, and continuity machinery as transcribe; text_only mode swaps the prompt (POLISH_PROMPT), skips image preparation, and disables citation search.

Multi-input behavior

Command Default With --merge
youtube URL1 URL2 Independent (2 outputs) — (no merge; one note per video)
local f1.mp4 f2.mp4 Independent (2 outputs) Merged (1 output)
transcribe f1.md f2.md Independent (2 outputs) Merged (1 output)
polish f1.md f2.md Each updated in place — (polish never merges)

--merge exists for stitching one long event (e.g., a workshop recorded as several local files) into a single note. Without it, transcribe -o takes a directory when there are several inputs (a file path or -t needs a single output, so it is a usage error). A merged output keeps each source file as its own section: an H1 heading per original file (its title), with H2 timestamp headings restarting under each. The overall generated title lives in the frontmatter only. Parts are processed in separate batches with continuity context reset at each boundary, so transcription never bleeds across recordings.

About

Video frame capture and AI transcription pipeline for markdown notes

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Contributors

Languages