[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"project-95906":3},{"id":4,"name":5,"fullName":6,"owner":7,"repo":5,"description":8,"homepage":8,"htmlUrl":8,"language":9,"languages":8,"totalLinesOfCode":8,"stars":10,"forks":11,"watchers":12,"openIssues":13,"contributorsCount":13,"subscribersCount":13,"size":13,"stars1d":13,"stars7d":13,"stars30d":14,"stars90d":13,"forks30d":13,"starsTrendScore":13,"compositeScore":15,"rankGlobal":8,"rankLanguage":8,"license":16,"archived":17,"fork":17,"defaultBranch":18,"hasWiki":19,"hasPages":17,"topics":20,"createdAt":8,"pushedAt":8,"updatedAt":21,"readmeContent":22,"aiSummary":23,"trendingCount":13,"starSnapshotCount":13,"syncStatus":24,"lastSyncTime":25,"discoverSource":26},95906,"ffmpeg-skill","kajisho5\u002Fffmpeg-skill","kajisho5",null,"Python",815,53,1,0,558,9.2,"MIT License",false,"main",true,[],"2026-09-21 02:04:29","\u003Cp align=\"center\">\n  \u003Cimg src=\"assets\u002Flogo.png\" alt=\"FFmpeg Skill: media processing for AI agents\" width=\"760\">\n\u003C\u002Fp>\n\n\u003Ch1 align=\"center\">ffmpeg-skill\u003C\u002Fh1>\n\n\u003Cp align=\"center\">\u003Cstrong>Give your coding agent a video editor.\u003C\u002Fstrong>\u003C\u002Fp>\n\n\u003Cp align=\"center\">\n  Local FFmpeg · No cloud · No API keys · Python standard library\u003Cbr>\n  Claude Code · Cursor · Codex · MCP\n\u003C\u002Fp>\n\n\u003Cp align=\"center\">\n  \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fkajisho5\u002Fffmpeg-skill\u002Factions\u002Fworkflows\u002Fci.yml\">\u003Cimg src=\"https:\u002F\u002Fgithub.com\u002Fkajisho5\u002Fffmpeg-skill\u002Factions\u002Fworkflows\u002Fci.yml\u002Fbadge.svg\" alt=\"tests\">\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fkajisho5\u002Fffmpeg-skill\u002Factions\u002Fworkflows\u002Fcodeql.yml\">\u003Cimg src=\"https:\u002F\u002Fgithub.com\u002Fkajisho5\u002Fffmpeg-skill\u002Factions\u002Fworkflows\u002Fcodeql.yml\u002Fbadge.svg\" alt=\"CodeQL\">\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fwww.npmjs.com\u002Fpackage\u002Fffmpeg-skill\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fnpm\u002Fv\u002Fffmpeg-skill\" alt=\"npm\">\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fwww.npmjs.com\u002Fpackage\u002Fffmpeg-skill\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fnpm\u002Fdm\u002Fffmpeg-skill\" alt=\"npm downloads\">\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fkajisho5\u002Fffmpeg-skill\u002Fstargazers\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fgithub\u002Fstars\u002Fkajisho5\u002Fffmpeg-skill\" alt=\"GitHub stars\">\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fkajisho5\u002Fffmpeg-skill\u002Fcommits\u002Fmain\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fgithub\u002Flast-commit\u002Fkajisho5\u002Fffmpeg-skill\" alt=\"last commit\">\u003C\u002Fa>\n  \u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fpython-3.9%20%7C%203.13-blue\" alt=\"Python 3.9 and 3.13 tested\">\n  \u003Ca href=\"#ffmpeg-compatibility\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fffmpeg-5.1%20%7C%206.1%20%7C%207.1%20%7C%208%20%7C%209%20tested-orange\" alt=\"FFmpeg 5.1, 6.1, 7.1, 8 and 9 tested in CI\">\u003C\u002Fa>\n  \u003Ca href=\"LICENSE\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Flicense-MIT-green\" alt=\"MIT\">\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fsponsors\u002Fkajisho5\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fsponsor-%E2%9D%A4-ea4aaa?logo=githubsponsors\" alt=\"Sponsor\">\u003C\u002Fa>\n\u003C\u002Fp>\n\n```bash\nnpx ffmpeg-skill\n```\n\n\u003Ctable>\n  \u003Ctr>\n    \u003Ctd width=\"50%\">\u003Cimg src=\"docs\u002Fdemos\u002Fcaptions_pop_karaoke.gif\" alt=\"animated captions with a karaoke highlight\">\u003Cbr>\u003Csub>\u003Ccode>caption.py --animate pop --karaoke\u003C\u002Fcode>\u003C\u002Fsub>\u003C\u002Ftd>\n    \u003Ctd width=\"50%\">\u003Cimg src=\"docs\u002Fdemos\u002Freframe_crop.gif\" alt=\"16:9 reframed to 9:16\">\u003Cbr>\u003Csub>\u003Ccode>fit.py --aspect 9:16 --fit crop\u003C\u002Fcode>\u003C\u002Fsub>\u003C\u002Ftd>\n  \u003C\u002Ftr>\n  \u003Ctr>\n    \u003Ctd width=\"50%\">\u003Cimg src=\"docs\u002Fdemos\u002Fsilence_removal.gif\" alt=\"silence removed, timeline shorter\">\u003Cbr>\u003Csub>\u003Ccode>silence.py --threshold -35\u003C\u002Fcode>\u003C\u002Fsub>\u003C\u002Ftd>\n    \u003Ctd width=\"50%\">\u003Cimg src=\"docs\u002Fdemos\u002Floudness.gif\" alt=\"waveform before and after loudness normalisation\">\u003Cbr>\u003Csub>\u003Ccode>loudness.py -I -14 --tp -1\u003C\u002Fcode>\u003C\u002Fsub>\u003C\u002Ftd>\n  \u003C\u002Ftr>\n\u003C\u002Ftable>\n\nLeft half is the input, right half is what the command produced. **[All 53 before\u002Fafter demos, with the exact command under each one →](docs\u002Fdemos.md)** — all of it generated from synthetic footage by `python3 demos\u002Fbuild.py`, so you can rebuild every frame of it yourself.\n\n`ffmpeg-skill` is an [Agent Skill](https:\u002F\u002Fdocs.anthropic.com\u002Fen\u002Fdocs\u002Fagents-and-tools\u002Fagent-skills) for Claude Code, Cursor, Codex and any agent that reads `SKILL.md`. It teaches the agent a fixed workflow (probe → edit losslessly where possible → check → verify) and ships **42 tools** that do the actual work with `ffmpeg` \u002F `ffprobe`: cut, join, silence removal, fit to duration and aspect, captions and karaoke, overlays and motion graphics, HDR → SDR and LUTs, audio clean-up and typed dynamics, sync with drift correction, multicam, loudness, delivery checks, whole-edit project rendering, batch folders. Every tool is also callable as an MCP tool (`tools\u002Flist` advertises a core 12 by default to keep client context small, with `FFMPEG_SKILL_MCP_FULL=1` listing all 42; every tool is reachable by name through `tools\u002Fcall` either way), and the whole set is described by a machine-readable contract.\n\nIf `ffmpeg` and `python3` are on your PATH, it works: offline, on footage you would rather not upload.\n\n> **SPEC** (Self-Producing Execution Contract): each tool's `input_schema` — the part of its\n> contract and MCP tool definition that has to track the CLI flag-for-flag — is never\n> hand-authored beside the code. It is derived, at run time, from the same `argparse` parser that\n> already defines the CLI, and CI fails the build if any of it drifts.\n> → [full explanation](#what-is-spec)\n\n---\n\n## Standalone, and in an ecosystem\n\n**Standalone**, this is a local FFmpeg engine: probe → edit → verify, `npx ffmpeg-skill` and nothing else. No API key, no account, no other repo required. Everything above and below this section describes that standalone tool, and none of it changes if you never read the rest of this one.\n\n**In [kajisho5](https:\u002F\u002Fgithub.com\u002Fkajisho5)'s wider video-production ecosystem**, this repo is the *hands*: it cuts, measures and exports files, and reports back in structured JSON. It does not decide what to cut, whether a deliverable is approvable, what makes a highlight interesting, or what a caption should say (the user's cue text is burned as written, never rewritten to fit) — those are a *brain*'s job, sitting in front of this engine, not inside it.\n\n| You want to... | Use |\n|---|---|\n| Cut \u002F join \u002F measure \u002F export a file right now | **this repo** (`ffmpeg-skill`), standalone |\n| Decide cut points, approve a deliverable, plan a whole edit | [`video-production-agent`](https:\u002F\u002Fgithub.com\u002Fkajisho5\u002Fvideo-production-agent) \u002F [`AI-video-production-OS`](https:\u002F\u002Fgithub.com\u002Fkajisho5\u002FAI-video-production-OS) |\n| Build a typed editing graph across a workspace, without writing raw `ffmpeg` | [`video-editing-skill`](https:\u002F\u002Fgithub.com\u002Fkajisho5\u002Fvideo-editing-skill) \u002F [`audio-production-skill`](https:\u002F\u002Fgithub.com\u002Fkajisho5\u002Faudio-production-skill) |\n\nOther repos in the ecosystem — [`media-analysis-skill`](https:\u002F\u002Fgithub.com\u002Fkajisho5\u002Fmedia-analysis-skill), [`transcription-skill`](https:\u002F\u002Fgithub.com\u002Fkajisho5\u002Ftranscription-skill), [`subtitle-skill`](https:\u002F\u002Fgithub.com\u002Fkajisho5\u002Fsubtitle-skill), [`thumbnail-skill`](https:\u002F\u002Fgithub.com\u002Fkajisho5\u002Fthumbnail-skill), [`color-grading-skill`](https:\u002F\u002Fgithub.com\u002Fkajisho5\u002Fcolor-grading-skill), [`motion-graphics-skill`](https:\u002F\u002Fgithub.com\u002Fkajisho5\u002Fmotion-graphics-skill), [`qc-skill`](https:\u002F\u002Fgithub.com\u002Fkajisho5\u002Fqc-skill) — read this repo's `contract --json`, its tools' `--json` output and `doctor`, the same way any agent framework would; this repo does not call into any of them. The dependency runs one way.\n\n---\n\n**Contents**\n[Standalone, and in an ecosystem](#standalone-and-in-an-ecosystem) · [Why](#why) · [Quick start](#quick-start) · [How it works](#how-it-works) · [Design principles](#design-principles) · [Tools](#tools) · [Audio](#audio-is-a-first-class-input) · [Built for agents](#built-for-agents) · [FFmpeg compatibility](#ffmpeg-compatibility) · [Tested on real footage](#tested-on-real-footage) · [Install](#install) · [Requirements](#requirements) · [Development](#development) · [Docs](#docs)\n\n---\n\n## Why\n\nAn agent that \"knows FFmpeg\" still guesses: it assumes a frame rate, picks a codec the container cannot hold, re-encodes a file that only needed a stream copy, and reports \"done\" without opening the result. This skill takes the guessing out:\n\n- **Real files first.** Every job starts with `probe.py`; the agent decides from the measured duration, fps, resolution, colour and audio layout, not from the file name.\n- **Structured tools, not shell strings.** Each operation is a script with typed arguments. Nothing runs through a shell; no filter graph is accepted from the caller.\n- **A contract the agent can read.** `contract --json` states, for every tool, what it takes, what it writes, which FFmpeg components it needs and how the result is verified. The MCP surface is derived from it.\n- **Verification after execution.** The result is probed, checked against the destination's spec and, when the picture changed, looked at as a contact sheet.\n- **Local first.** No cloud, no API keys, no Python dependencies. Optional local transcription is used when a whisper is installed, never required — by `caption.py --transcribe` and, since 1.17, by `silence.py --filler --transcribe`; both take a transcript you already have instead, and both refuse with the install lines rather than guessing.\n\n## Quick start\n\n```bash\n# 1. install the skill for Claude Code (Cursor: --cursor, Codex: --codex, all three: --all)\nnpx ffmpeg-skill\n\n# 2. check the machine: ffmpeg, ffprobe and every FFmpeg component the tools need\nnpx ffmpeg-skill doctor\n\n# 3. (for agent frameworks) read the machine-readable contract\nnpx ffmpeg-skill contract --json | head -40\n```\n\nAlready installed? re-run `npx ffmpeg-skill` to refresh `~\u002F.claude\u002Fskills\u002Fffmpeg-skill`. Copies are not updated automatically.\n\n`doctor`'s overall `ok` and a single tool's `usable: no` are different signals: `ok` means nothing *required by every tool* is missing, but a plain Homebrew `ffmpeg` on macOS can still be `ok` while `caption.py` specifically can't run (no `subtitles` filter) — check `doctor --json`'s `tools` field for the per-tool answer, not just `ok`.\n\nThen talk to your agent:\n\n> \"Take `interview.mp4`, keep 0:45–3:10 and 5:00–6:30, and make it exactly 60 seconds for Reels.\"\n\nThe agent runs `probe.py`, `cut.py --segments 0:45-3:10,5:00-6:30`, `fit.py --duration 60 --aspect 9:16 --fit crop`, `export.py --preset reels`, `check.py --platform reels` and `look.py`, then reports \"final.mp4: 59.98 s, 1080×1920, 30 fps, AAC stereo\" with the contact sheet it inspected.\n\n### Deliver to a platform\n\nOne command per destination, with the app's own UI taken into account:\n\n```bash\npython3 $S\u002Frender.py talk.mp4 --template tiktok --cues cues.txt\n```\n\nThat fills the shipped `templates\u002Ftiktok.json`: 9:16 crop, captions popped word by word *above*\nTikTok's description bar and clear of its like column, −14 LUFS, the `tiktok` export preset, and\na `check.py --platform tiktok` on the file it wrote. The caption size comes from the delivery\ntable, so since 1.17.1 the filled project also states `\"fit_size\": \"on\"`: a size nobody asked for\nshrinks to fit the cue instead of splitting the sentence across two cues. A template that states\nits own `fit_size`, and a `--brand` that states a caption size, still win, and a project with\n`\"fit_size\": \"off\"` renders the captions 1.17.0 rendered. Templates ship for `tiktok`, `reels`,\n`shorts`, `youtube-shorts`, `youtube`, `x`, `linkedin`, `facebook` and `podcast`;\n`--template all` (or a comma-separated list) renders every destination from the same edit and\nwrites a `\u003Cname>_pack.md` table of what each one produced. Files land next to the input unless\n`-o` says otherwise, and `--dry-run` shows every planned command rather than a result.\n`render.py --list-templates` prints them with their frames, limits and safe zones. Alias\nspellings work everywhere a platform is named (`youtube-shorts` = `shorts`, `ig` = `reels`,\n`twitter` = `x`, `fb` = `facebook`).\n\nThe tools also work on their own, from any shell:\n\n```bash\nS=~\u002F.claude\u002Fskills\u002Fffmpeg-skill\u002Fscripts\npython3 $S\u002Fprobe.py input.mp4 --compact\npython3 $S\u002Ffit.py input.mp4 --duration 60 --aspect 9:16 --dry-run    # print the plan, run nothing\npython3 $S\u002Fexport.py input.mp4 --preset reels --json                 # structured result with a probe of the output\n```\n\nOn Windows in Git Bash, `python3` is only on PATH if Python was installed from the Microsoft Store; a python.org install exposes `python` (or the `py` launcher) instead — replace `python3` with `python` above if you see a \"command not found\". `bin\u002Finstall.js` and `doctor`\u002F`contract` already handle this for you; only the raw script examples above need it spelled out manually.\n\nMore requests and the commands behind them: [examples\u002FREADME.md](examples\u002FREADME.md). To see everything run end-to-end on generated footage: `npm run demo` (the gallery in [docs\u002Fdemos.md](docs\u002Fdemos.md)).\n\n## How it works\n\n```mermaid\nflowchart TD\n    U[User request] --> A[AI agent\u003Cbr\u002F>Claude Code · Cursor · Codex]\n    A -->|reads| S[SKILL.md\u003Cbr\u002F>workflow, request → tool map, report format]\n    A -->|runs| T[Structured tool\u003Cbr\u002F>scripts\u002F&lt;name&gt;.py, typed argparse flags]\n    T --> C[Contract\u003Cbr\u002F>input schema · role · capabilities · verification policy]\n    C --> D[Capability detection\u003Cbr\u002F>doctor: available \u002F missing \u002F unknown]\n    D --> F[FFmpeg execution\u003Cbr\u002F>no shell, stream copy when possible]\n    F --> V[Verification\u003Cbr\u002F>probe · check · look.py contact sheet]\n    V --> R[Structured result\u003Cbr\u002F>--json: status, output, commands, probe]\n    R --> A\n```\n\nOver MCP the same tools are reached through a transport that holds no tool table of its own:\n\n```mermaid\nflowchart LR\n    M[MCP client\u003Cbr\u002F>Claude Desktop · Cursor · any client] --> P[mcp\u002Fserver.py\u003Cbr\u002F>stdio JSON-RPC]\n    P -->|tools\u002Flist| C[Contract-derived ToolSpecs\u003Cbr\u002F>names · order · inputSchema]\n    P -->|tools\u002Fcall| T[scripts\u002F&lt;name&gt;.py]\n    C -.derived from.-> K[scripts\u002F_contract.py]\n    T -.described by.-> K\n```\n\nNames, order and `inputSchema` in `tools\u002Flist` are translated from each tool's argparse parser at start-up, so a new flag or a new script appears in MCP with no edit to `mcp\u002F`. A test copies the skill, adds, removes and edits a script, and reads `tools\u002Flist` again to prove it.\n\n## Design principles\n\nThese are the rules the skill file gives the agent and the code enforces.\n\n1. **Probe first.** No tool decides from the file name. `probe.py` measures duration, fps (with variable-frame-rate detection), resolution, rotation, bit depth, HDR format including Dolby Vision, colour tags and every audio stream before anything is cut.\n2. **Lossless when possible.** `cut.py`, `join.py` and `loudness.py` stream-copy what they do not need to touch. Re-encoding happens only when it must: frame-accurate cuts, filters, format changes, or a keyframe farther than the tolerance.\n3. **Plan before render.** Every tool takes `--dry-run` (print the ffmpeg command lines, write nothing), `--plan FILE` (the dry run saved as a plan with fingerprinted inputs that `render.py FILE` executes later, refusing if an input changed), `--json` (structured result with a probe of the output), `--fast` (preview quality), `--progress` (percent and ETA), `--timeout` (a hung ffmpeg is killed and reported, never waited on forever; Ctrl-C or SIGTERM likewise stops the running ffmpeg, removes its partial output and reports `kind: interrupted`) and `--overwrite` (explicit consent before an existing output is replaced). A test runs every tool under `--dry-run` behind a fake ffmpeg and asserts that no ffmpeg call happened and no file appeared.\n4. **Machine-readable contract.** `contract --json` describes all 42 tools: input schema generated from the parser, output schema, role, required and conditional FFmpeg capabilities, dry-run support, the verification tools to run afterwards, whether a visual check is required, `mutates_input: false`. `provides` lists all 42 by a cross-repository Capability id (`ffmpeg-skill.cut`, `ffmpeg-skill.loudness`, ...) for [`kajisho5\u002FAI-video-production-OS`](https:\u002F\u002Fgithub.com\u002Fkajisho5\u002FAI-video-production-OS)'s `CapabilityContract.provides` — see `docs\u002Fcontract.md`.\n5. **Contract-derived MCP.** `mcp\u002Fserver.py` builds its `tools\u002Flist` from the contract. Tool names, order and `inputSchema` cannot drift from the scripts; a test keeps the two byte-identical.\n6. **Capability detection.** `doctor` reads `ffmpeg -encoders \u002F -filters \u002F -bsfs` and reports which of the components the tools need are present on this build (libx264, libass, zscale, loudnorm, xfade, …), before a job fails inside ffmpeg.\n7. **Unknown is not missing.** When a listing cannot be read (a layout the parser does not know, ffmpeg exiting non-zero) the affected capabilities are `unknown`: never `missing`, never silently `available`. An installed filter is not reported absent; a failed detection is not a pass.\n8. **Verify the result.** The output is probed, and when the picture changed (captions, overlays, crops, colour, transitions) the agent runs `look.py` and inspects the PNG. The report is not finished until its `Look:` line names that image; audio-only jobs say `Look: not needed`. **\"Inspects\" means the calling agent's own vision, not a feature of this skill:** `look.py` only renders a PNG; nothing in this repository detects faces, products, subjects, or \"the interesting part\" of a frame or a scene. When a crop or reframe needs to keep a specific part of the frame (`fit.py --fit crop --crop-x\u002F-y`, see [Tools](#tools)), it is the multimodal agent looking at that PNG and choosing the anchor — a non-visual caller (a script, a CLI user without eyes on the sheet) has to supply that decision itself, and the default is a plain centre crop. Likewise `scenes.py --highlights` ranks candidate scenes by a measured proxy (`--rank-by audio` or `--rank-by duration`), never by content; it is the agent that turns a look at the sheet into a judgement.\n9. **One label per report.** A finished job is `Done:`, a failure or a refusal is `Failed:`, and a partial result is `Done:` with the shortfall named in `Notes:` — never a third label such as `Done (partially):`.\n10. **Keep originals.** No tool overwrites its input. Outputs are new files named `\u003Cinput>_\u003Coperation>.\u003Cext>` unless told otherwise, and a test hashes every input after the run.\n\n## Tools\n\n42 public tools, all Python 3.9 standard library, all with `--help`, `--dry-run`, `--json`, `--plan FILE` (a dry run written as a plan `render.py` executes later), non-zero exit and a reason on stderr on failure. `--json-brief` (1.11.0) prints the same result document trimmed to what a caller acts on — status, output, `verified`, a compact `summary` of the output probe, the tool's own keys and the command count instead of the command lines — for roughly a third of the bytes; `--json` itself is unchanged. Every re-encoding tool takes `--codec h264|hevc|av1|prores` and `--quality N` (1.8), and every time flag takes seconds, `mm:ss`, `hh:mm:ss.fff` or SMPTE `hh:mm:ss:ff` with an optional `@fps` suffix (1.9).\n\n**Analysis and inspection**\n\n| Tool | What it does |\n|---|---|\n| `probe.py` | Duration, fps (+ VFR detection), resolution, codecs, bit depth, HDR format incl. Dolby Vision (`hdr` for BT.2020 or PQ\u002FHLG, `hdr_signal` for a real PQ\u002FHLG\u002FDV transfer only), colour space, rotation, every audio stream; `--analyze` flags Log footage |\n| `scenes.py` | Scene changes, audio peaks, highlight proposals (`--rank-by audio` loudest, or `--rank-by duration` longest — both proxies, not \"best\") and a per-scene sheet; cut list for `cut.py --segments`; `--beats` measures the music's beat grid (tempo, beat times, confidence); `--shots` classifies each shot static\u002Fpan\u002Fmotion by measured optical flow; `--audio-peaks` lists per-second dBFS; `--speech` reports a speech-vs-music ratio, not a classification |\n| `look.py` | Contact sheet, single frames, side-by-side comparison as PNG so the agent can see what it made; `--safe NAME` shades the zones a platform's own UI covers |\n\n**Editing**\n\n| Tool | What it does |\n|---|---|\n| `cut.py` | In\u002Fout or multi-segment cuts, lossless `-c copy` first, re-encode fallback, `--accurate` for frame-exact video and sample-exact audio; reports `precision`; `--snap beats` moves the in\u002Fout points onto a measured beat, or refuses when there is no measurable pulse |\n| `join.py` | Concatenate clips with xfade transitions, normalising size, fps, sample rate and channel layout (the widest clip's, or `--channels`); audio-only inputs are joined as audio |\n| `silence.py` | Detect and remove dead air (jump cuts) with a margin around speech; list or export the cut list; `--filler` also removes filler words, but only where a speech engine timed them; `--speech-aware` keeps a breath inside a sentence and cuts only at sentence-boundary pauses, composing with `--filler` through the same `keep_ranges()` |\n| `fit.py` | Fit to a duration (pitch-preserving speed change or trim, smooth slow-mo) and\u002For aspect ratio (pad, crop or `--fit blur`'s blurred, darkened fill, with `--crop-x`\u002F`--crop-y` to keep an off-centre subject) and\u002For exact `--width`\u002F`--height`; rotate 90\u002F180\u002F270, flip h\u002Fv; force constant fps |\n| `crop.py` | Crop to an exact pixel rectangle (`--x --y --width --height`) — distinct from `fit.py --fit crop`, which crops to an aspect ratio it computes itself |\n| `cropdetect.py` | Measure existing black letterbox\u002Fpillarbox bars and report the `crop.py`-ready rectangle that removes them — analysis only, writes no file; `--motion-centre` reports the per-second motion centroid instead, for a 9:16 reframe — report only, the calling agent picks the crop, no auto-reframing |\n| `deinterlace.py` | Deinterlace interlaced source footage (`yadif`), `--mode frame`\u002F`field`, `--parity` |\n| `denoise.py` | Reduce video noise\u002Fgrain (`hqdn3d`), `--strength low\u002Fmedium\u002Fhigh` or individual spatial\u002Ftemporal overrides |\n| `redact.py` | Blur or pixelate an exact pixel rectangle for the whole clip (privacy\u002Fcompliance redaction) |\n| `sphere.py` | Extract a flat rectilinear viewport from a 360\u002Fspherical video (`--yaw --pitch --roll --h-fov --v-fov`); no subject tracking, only the aim you give it |\n| `straighten.py` | Rotate by an arbitrary angle for horizon correction (`--degrees`, `--fit crop\u002Fpad`) — distinct from `fit.py --rotate`'s exact 90-degree turns |\n| `insert.py` | Turn a still image into a silent, fixed-duration video clip (title card, end slate) at an exact frame size \u002F fps, with an optional Ken Burns zoom\u002Fpan |\n| `background.py` | Generate a solid-colour or two-colour gradient clip at an exact size\u002Fduration — no input file |\n| `reverse.py` | Reverse playback (video and, unless `--no-audio`, audio) |\n| `stabilize.py` | Two-pass motion stabilisation (`vidstabdetect`\u002F`vidstabtransform`) |\n| `sequence.py` | Numbered (`frame_%04d.png`) or glob-matched still images into a video |\n| `waveform.py` | Render an audio track as a waveform or spectrum visualization video (`showwaves`\u002F`showspectrum`) — for audio-only inputs with no picture worth showing; `--image` puts the visualisation over a still plate (an audiogram), with `--platform`, `--title` and burnt-in captions |\n| `freeze.py` | Hold a frame for N seconds (`--at`, `--hold`, `--mode insert\u002Fextend`) — an end-card hold or a comedic beat |\n| `pad.py` | Add black\u002Fsilent padding at the start and\u002For end of the timeline (`--start`, `--end`) — distinct from `fit.py --fit pad`'s per-frame letterbox bars |\n| `speedramp.py` | Step through different constant speeds across a clip via `--segment START-END:FACTOR` (repeatable) — distinct from `fit.py`'s single whole-clip speed factor |\n| `loop.py` | Repeat a clip `--times` N or to a target `--duration` — for background loops and filling a fixed slot length |\n| `broll.py` | Cut away to a B-roll clip over the A-roll for a window (`--insert B --at T --duration D`, repeatable) and come back; A's length and audio untouched by default |\n| `metadata.py` | Write container chapter markers from a `TIME TITLE` text file and title\u002Fartist\u002Fcomment tags, every stream copied bit for bit; `--auto-chapters` proposes the markers from measured pauses and scene cuts and names them `Chapter N` for you to rename |\n| `grid.py` | Composite `--cols`x`--rows` clips into one grid, each cell letterboxed and labelled with its filename by default (`--label none` to skip) |\n\n**Audio**\n\n| Tool | What it does |\n|---|---|\n| `audio.py` | Voice clean-up chain at three strengths (`--voice light\\|medium\\|strong`), FFT denoise, typed compressor \u002F limiter \u002F gate, music bed with sidechain ducking (`--duck-amount\u002F-threshold\u002F-attack\u002F-release`), a never-ducked effects bed (`--effects`), `--stereo-widen`, fades, 5.1 → stereo, track replacement, extraction (`-o out.wav`), `--audio-stream N` |\n| `sync.py` | Offset between two recordings by audio cross-correlation (1 ms, pure Python), clock-drift correction; aligned video or audio out (audio-to-audio only — no lip-sync\u002Fface detection); a third or later camera is an additional `more_sources` positional (`sync.py REF A B ...`), the `second` positional itself keeps its name |\n| `loudness.py` | Two-pass EBU R128 `loudnorm` to −14 LUFS \u002F −1 dBTP or any target (`--lra` for the range), video stream-copied; the written file is measured again and re-encoded until it meets `--tp` (lossy encoders overshoot); `--measure-only` |\n\n**Picture**\n\n| Tool | What it does |\n|---|---|\n| `caption.py` | Burn SRT\u002FASS with font, size, colour, outline, position; build SRT from timed plain text; wraps to the safe area — with `--platform`, the destination's own left\u002Fright safe zone, which since 1.17.2 is also what the burnt ASS states as its side margins — with a phrase-aware breaker (`--wrap phrase|measured`) and `--max-lines`\u002F`--min-duration`\u002F`--offset`; `--mode mux` takes a repeated `--srt file:lang` for several language-tagged, toggleable tracks in one file; picks a font by script for non-Latin text (`--lang`); animated and word-by-word karaoke timed to the speech energy or real word timings; `--fit-size` shrinks the size until a cue fits `--max-lines` instead of splitting the sentence (on by default on the delivery-template path since 1.17.1, where the size comes from the platform table); says `caption text unchanged` when it burned the cues exactly as given; optional local transcription |\n| `overlay.py` | Logos, watermarks and titles with position, time range, opacity, fades; `--platform NAME` keeps them clear of that destination's UI; `--video` for picture-in-picture, `--chromakey` for green-screen compositing |\n| `graphics.py` | Lower-thirds, title cards, chapter chips, progress bars, countdowns, corner bugs, social stickers, opening hook cards and meme captions drawn by FFmpeg from a brand kit; `--platform NAME` keeps them inside that destination's safe zone; `--text-render` routes shaping scripts through libass and `--emoji-assets` composites colour emoji |\n| `color.py` | HDR10 \u002F HLG \u002F Dolby Vision → SDR BT.709 tone mapping, DV layer stripping, 3D LUT (.cube), colour-tag rewriting, typed primary correction (exposure\u002Fcontrast\u002Fsaturation\u002Fgamma\u002Fwhite balance\u002Flift-gain\u002Flevels\u002Fcurves) |\n\n**Delivery**\n\n| Tool | What it does |\n|---|---|\n| `export.py` | Presets `youtube`, `youtube4k`, `reels`, `tiktok`, `shorts`, `linkedin`, `facebook`, `x`, `youtube-hdr` (HEVC Main10, source HDR tags kept), `youtube-av1`, `prores`, `h265`, `gif`, `copy`, all tagged BT.709 unless they carry HDR; `--normalize` meets the platform's loudness in the same call (`render.py` turns it on by default for platform presets) |\n| `proxy.py` | Small, low-bitrate proxy for downstream AI analysis\u002Fpreview\u002Fediting decisions — resize by `--width`\u002F`--scale`, proxy-grade `--crf` (deprecated alias of `--quality`), `--fps`, `--no-audio`; not a delivery preset |\n| `check.py` | PASS \u002F WARN \u002F FAIL against YouTube, Shorts, Reels, TikTok, X, LinkedIn, Facebook, broadcast and podcast specs, from the same delivery table the export presets and templates read (podcast also reports chapter markers and channel count), with the fix for each failure and a `format` \u002F `judgement` kind per row |\n| `report.py` | Single-file HTML delivery report: before\u002Fafter sheets, media facts, loudness, compliance, the commands run; `--pack` renders a social pack table |\n\n**Orchestration**\n\n| Tool | What it does |\n|---|---|\n| `render.py` | Render a whole edit from a declarative `project.json` (clips, transitions, captions, overlays including the social sticker\u002Fhook\u002Fmeme graphics, music and stem levels, loudness, export, chapter markers, check); `--init`, `--dry-run`, `--stop-after`, `--cache DIR`\u002F`--from STAGE` (reuse the stages that did not change); the captions block takes the fit-size policy (`fit_size`, `min_size`, `fit_size_scope`) and the run reports the caption stage's counts as `caption`; `--template NAME INPUT` renders a shipped delivery template (`--template all` writes the whole social pack plus its table) |\n| `batch.py` | Apply a step recipe or a project to a folder with a content-hash cache; `--watch`; `--jobs N` processes several files at once under one shared `--timeout` |\n| `multicam.py` | Align any number of cameras and recorders by audio (with drift correction) and cut between them from a switch list; `--switch energy` auto-cuts to whichever camera is loudest\u002Ftalking instead of a hand-built list, with `--min-shot` (minimum shot length) and `--edl` (write the cut list) |\n| `verify.py` | Run the toolchain on real device files and report PASS \u002F FAIL per step |\n\nNot tools, but part of the surface: `mcp\u002Fserver.py` (the MCP transport) and `scripts\u002F_contract.py` (`contract --json`, `doctor`). Per-flag reference for every tool: [references\u002Fscripts.md](references\u002Fscripts.md).\n\n## Audio is a first-class input\n\nWAV, FLAC, MP3, M4A\u002FAAC, OGG and Opus go through `probe`, `cut`, `join`, `silence`, `loudness`, `audio`, `sync` and `check --platform podcast` with the same commands as video. The output extension picks the codec: `-o out.wav` writes PCM, `-o out.flac` FLAC, `-o out.mp3` MP3, `-o out.m4a` AAC.\n\n- **Extraction.** An audio extension on a video input drops the picture: `audio.py talk.mp4 -o talk.wav`, or `--voice -o talk.m4a` to clean it on the way. `--audio-stream N` picks a track; `probe` lists them under `audio_streams`.\n- **Join.** `join.py intro.wav episode.m4a outro.wav -o full.flac` resamples every clip to one rate and channel layout and crossfades them (`--transition none` for a butt join). Audio and video inputs cannot be mixed in one join.\n- **Sample-accurate trims.** `cut.py talk.wav --start 1.2345 --end 2.3456 --accurate` trims at the sample; the JSON reports `precision` (`packet` for a stream copy, `sample` for PCM \u002F FLAC, `codec_frame` when a lossy encoder frames the audio again, `frame` for video) and the measured `duration_error_ms`. A `.wav` never receives compressed packets.\n- **Typed dynamics.** `audio.py --compress --comp-threshold -20 --comp-ratio 4`, `--limit --limit-ceiling -1`, `--gate --gate-threshold -45`. Each flag is one documented option of FFmpeg's `acompressor`, `alimiter` or `agate`, range-checked before ffmpeg runs; no filter string is accepted from the caller.\n- **Loudness.** `loudness.py talk.wav -I -16 --tp -1.5 -o talk.m4a` for podcast levels; `check.py talk.m4a --platform podcast` measures LUFS and true peak and reports the chapter markers and channel count.\n\nPicture tools (`fit`, `caption`, `overlay`, `graphics`, `color`, `export`, `scenes`, `look`) refuse an audio file with \"input has no video stream\" instead of inventing a picture.\n\n## Built for agents\n\n### What is SPEC?\n\n**SPEC** (Self-Producing Execution Contract) is the name this project's author,\n[kajisho5](https:\u002F\u002Fgithub.com\u002Fkajisho5), gave the pattern the tool layer is built on: each tool's\n`input_schema` — the part of its contract that has to track the CLI exactly, flag for flag — is\nnever hand-authored side by side with the code. It is derived, at run time, from the one thing\nthat has to be correct for the CLI to work at all: the script's own `argparse` parser.\n\nConcretely, `scripts\u002F_contract.py`'s `_capture_parser()` imports every tool script and\nintercepts its `parse_args()` call to get the live, fully-built parser object — flags, types,\nchoices, defaults, required\u002Fpositional, mutually exclusive groups, all of it. `input_schema` is\nbuilt straight from that object. (The rest of a `ToolSpec` — `role`, `capabilities`, `inputs`,\n`outputs`, `output_schema` — comes from a hand-authored table, `TOOL_META`, since those facts\naren't things a parser can express; only `input_schema` is parser-derived.)\n\n- **The contract**'s `input_schema` for every tool is generated from the live parser directly.\n- **SKILL.md is two-tier** (1.11.0): the file the agent loads every session keeps the workflow, the\n  request→script table and one line per gotcha; the long-form detail lives in `references\u002Fgotchas.md`\n  and the other `references\u002F` files, read only when a job needs it.\n- **The MCP server** (`mcp\u002Fserver.py`) carries no schema of its own; `tools\u002Flist` is translated\n  straight from the contract, `input_schema` included.\n- **The docs** (`docs\u002Fcontract.md`'s field reference, this README's tool table) describe the same\n  shape. `tests\u002Ftest_contract.py` runs on every CI run and fails the build if any of them drift\n  out of sync with what the code actually does — it catches drift, it doesn't fix it for you.\n\nSo adding a flag to a script's `argparse` block updates `input_schema` and the MCP tool\ndefinition with no second edit, and a docs page or `TOOL_META` entry that falls behind fails CI\nrather than drifting silently. There is no separate `input_schema` file to forget to update.\n\n### Machine-readable contract\n\n```bash\nnpx ffmpeg-skill contract --json            # or: python3 scripts\u002F_contract.py --json\nnpx ffmpeg-skill contract --json --static   # without environment detection\n```\n\nThe contract is generated from the code that runs, not maintained beside it. For each of the 42 tools (`ffmpeg-skill\u002F\u003Cname>`) it states:\n\n| Field | Meaning |\n|---|---|\n| `input_schema` | generated from the tool's argparse parser: properties, types, enums, defaults, required, positional order, mutually exclusive groups |\n| `output_schema` | what `--json` prints: `status`, `output`, `commands`, `probe`, plus tool-specific fields (`precision`, `checks`, `offset_seconds`, …) |\n| `role` | `analysis`, `analysis_and_execution`, `execution` or `verification` |\n| `capabilities` | the FFmpeg encoders, filters and bitstream filters the tool always needs, and the ones needed only for a flag or input |\n| `supports_dry_run`, `supports_json`, `supports_json_brief` | measured by the tests, not declared |\n| `verification` | which tools to run on the output afterwards (`probe`, `check`, `look`) |\n| `requires_visual_verification` | the picture changed; inspect the contact sheet |\n| `audio_only`, `video_required` | whether an audio-only input is accepted or refused |\n| `mutates_input` | always `false` |\n| `idempotency_hint` | `bit_exact`, `content_equivalent`, `cached` or `environment_dependent` |\n\nNext to the tool list the document carries a top-level `deprecated` list (1.10): what 2.0.0 removes, since when, the replacement and the surface it lives on. `docs\u002Fcontract.md` \"What 2.0 changes\" is written from it.\n\n`contract_version` (1.0) is separate from the skill version, so a consumer can pin the shape and read the version for provenance. The document also states the invocation mapping (structured arguments → argv), the JSON shapes for success and failure (`{\"status\": \"failed\", \"error\": {\"kind\": \"input | ffmpeg | output | missing_tool | timeout | verification | interrupted\", \"message\": …}}`), and that no tool runs a shell or executes anything other than the named script, `ffmpeg` and `ffprobe`. Field-by-field reference: [docs\u002Fcontract.md](docs\u002Fcontract.md).\n\n### MCP\n\n```json\n{\"mcpServers\": {\"ffmpeg-skill\": {\"command\": \"python3\", \"args\": [\"\u002FUsers\u002Fyou\u002F.claude\u002Fskills\u002Fffmpeg-skill\u002Fmcp\u002Fserver.py\"]}}}\n```\n\nOn Windows, `python3` is only on PATH if Python was installed from the Microsoft Store; a python.org install exposes `python` (or the `py` launcher) instead — if your MCP client reports the server failed to start, change `\"command\"` above to `\"python\"` (or the full path from `where python`).\n\n`mcp\u002Fserver.py` is a stdio JSON-RPC transport with no tool table of its own. `tools\u002Flist` is derived from the contract at start-up, in contract order, with `inputSchema` translated from each tool's `input_schema`. By default it lists only the core 12 (`render`, `look`, `caption`, `export`, `check`, `fit`, `cut`, `audio`, `loudness`, `graphics`, `silence`, `probe`) so a client doesn't pay context for 30 schemas it rarely calls directly; set `FFMPEG_SKILL_MCP_FULL=1` to list all 42. Every tool, listed or not, is callable through `tools\u002Fcall`, which maps structured arguments to argv and runs the named script; a raw `argv` form is accepted for compatibility and marked non-canonical. `python3 mcp\u002Fserver.py --list` prints the tools; `--call probe '{\"inputs\": [\"a.mp4\"]}'` runs one from the shell.\n\n`FFMPEG_SKILL_MCP_LEAN=1` in the server's environment drops `json` and `progress` from every `inputSchema`: they are transport flags the server sets itself, not tool arguments, and 2.0 drops them unconditionally. It is opt-in, independent of `FFMPEG_SKILL_MCP_FULL`, and the default `tools\u002Flist` stays byte-identical (aside from the core-12 filter) to the CLI surface the contract promises.\n\n### Capability detection\n\n```bash\nnpx ffmpeg-skill doctor          # human-readable\nnpx ffmpeg-skill doctor --json   # available \u002F missing \u002F missing_optional \u002F unknown \u002F detection \u002F errors \u002F tools \u002F gpu_encoders\n```\n\n`doctor` reads `ffmpeg -encoders`, `-filters` and `-bsfs` and resolves every capability the contract declares against this machine's build. Three states per capability: `available`, `missing`, `unknown`. Exit 0 when everything required is available, 1 when something required is missing, 2 when nothing is proven missing but a required capability is unknown. With detection on (the default), `contract --json` carries the same lists under `capabilities`. `doctor` also reports fonts: the default drawtext family, and `fonts.scripts` — one `available`\u002F`missing`\u002F`unknown` per writing system (ja, zh, ko, ar, he, hi, th, ru, el) with the file it would use — so \"can this machine render Korean captions\" is answered before the job, not after. `doctor --json`'s `tools` field folds that down to one answer per tool — `{\"caption\": {\"usable\": \"no\", \"missing\": [\"filter:subtitles\"], \"fix\": \"...\"}, ...}` — so \"is `doctor` overall `ok`\" and \"can I run `caption.py` on this machine\" are answered separately: a plain Homebrew `ffmpeg` is `ok` for tools that don't need `subtitles`\u002F`drawtext`\u002F`zscale`, while `caption`'s own `usable` is `\"no\"`.\n\n`doctor --json`'s `gpu_encoders` reports which GPU-backed encoders (`nvenc`, `videotoolbox`, `qsv`, `vaapi`, `amf`) this ffmpeg *build* was compiled with — read from `-encoders` alone, so it proves the capability shipped, not that the GPU\u002Fdriver on this machine will actually accept a job (that needs a real encode, which `doctor`'s introspection never runs). No tool here uses one yet — every tool still assumes CPU x264\u002Fx265 — so this is purely informational and never affects `ok` or any tool's `usable`. GPU-accelerated encoding stays deliberately off the roadmap until there's a real-hardware-verified design for it (build-presence alone is not proof a job will succeed) — not a promised feature, just an honest \"not yet, and not without proof it actually works.\"\n\n## Gotchas and best practices\n\nThe short list for humans. The agent-facing version, with the reasoning, is the \"Things that look right but are wrong\" and \"Gotchas\" sections of [SKILL.md](SKILL.md).\n\n- **Variable frame rate (phone and screen recordings).** `probe.py` flags it; every re-encoding tool conforms to a constant rate automatically, and `cut.py` switches to frame-accurate mode on its own because copy-cuts on VFR land on the wrong frame. Choose the rate yourself with `fit.py input.mp4 --fps 30` when the measured average is odd.\n- **Lossless cuts snap to keyframes.** A stream-copy cut can start up to one GOP earlier than asked. `cut.py` re-encodes when the snap exceeds 0.5 s (`--tolerance` changes the limit). For a strictly lossless file pass `--tolerance -1`, and expect the cut to land on the nearest earlier keyframe; the JSON result lists them under `nearest_keyframes`.\n- **HDR stays HDR.** When the probe reports HDR (HDR10, HLG, Dolby Vision, BT.2020), the tools keep it rather than flatten it. Convert deliberately with `color.py --to-sdr` before H.264 deliverables or LUT work. `export.py` platform presets are SDR and warn on HDR input.\n- **Loudness targets.** −14 LUFS \u002F −1 dBTP for YouTube and social platforms (the `loudness.py` default), `-I -16 --tp -1.5` for podcasts, `-I -23` for broadcast. A clip measured at −40 LUFS or below is room tone, not content; raising it raises the noise. Check true peak as well as LUFS: `check.py file --platform podcast` measures both.\n- **Frame changes first, text second.** Captions and overlays burned before a crop or resize end up off-frame. Reframe, then caption.\n- **Cropping 16:9 to 9:16 discards 70 % of the width.** `fit.py --fit crop` centres by default; pass `--crop-x`\u002F`--crop-y` toward the subject, or pad with `--fit pad --pad-fill blur`. Look at the contact sheet before deciding.\n- **Phrase-aware caption breaking (1.16).** `caption.py`\u002F`graphics.py --wrap phrase` (the default) never breaks inside a word or on the wrong side of a hyphen, never leaves a lone digit, kana or punctuation pair on a line, prefers Japanese sentence ends and particles over a mid-word break, and never ends a line on an article or preposition. All four are penalties over break positions that already fit, so no line is widened and the line count never changes; `--wrap measured` restores 1.15's width-only wrap. The text itself is never rewritten or shortened. Since 1.16.1 a Thai run and a katakana word are never broken inside (Thai writes no space inside a phrase and there is no dictionary: the break goes where you put a space or `|`), and a line wider than the safe width is reported as `overlong` with the fix named.\n- **Caption size fitted to the cue (1.17, reachable from the templates since 1.17.1).** At a platform caption size a line holds about six em, so an ordinary sentence needs four lines and `--max-lines 2` used to cut it into consecutive cues — half the sentence arriving late. `caption.py --fit-size` (default `auto`) walks the size down until every cue fits, *before* laying the cues out, with a legibility floor of 4.5 % of the frame height (`--min-size`, default 13 ASS units). An explicit `--size` or a `brand.json` size is a statement about the look and is never overridden — which in 1.17.0 also silenced the fitter on every `render.py --template` run, since a template fills the size from the platform table; 1.17.1 marks that size as the default it is (`\"fit_size\": \"on\"` in the filled project), so the type shrinks and no cue is split. `--fit-size off` restores that earlier behaviour byte for byte, and the caption text is still never rewritten to make it fit.\n- **Beat-synced cuts (1.17).** `scenes.py --beats` reports the measured grid — tempo, beat times, and a confidence built from how far the winning autocorrelation lag stands above the others and how many onsets land on it. `cut.py --snap beats` (and a `\"snap\"` block in a `render.py` project) moves in\u002Fout points to the nearest beat within `--snap-tolerance` — and only onto the grid points a measured onset actually marks, never onto the regular grid's continuation through a passage with no music in it. Below `--min-confidence` it **refuses**: a cut point may move to a measured beat and may not appear from one, so speech and ambience get an honest \"no steady pulse here\" instead of an invented grid.\n- **Filler words (1.17).** `silence.py --filler --words transcript.json` removes \"um\" and \"uh\" through the same cut graph the silences use. Never without measured word timings — there is no heuristic that finds an \"um\" without them that would not also cut real speech — and `like`, `tipo` and `cioè` are deliberately not in the built-in lists, because a discourse marker is a content word. Whisper stays optional: `--transcribe` with no engine installed refuses and names the three installs.\n- **Throughput (1.17).** `batch.py --jobs N` runs several files at once, capped at `min(N, cpu_count, 8)` and sharing one `--timeout` budget rather than one per item; the per-item table keeps its order. `render.py --cache DIR` reuses stages whose inputs and arguments did not change, so swapping an export preset re-runs export only. The cache is opt-in with no default directory, and the ffmpeg, skill and contract versions are part of every key, so a cache is never reused across them.\n- **Audiogram (1.16).** `waveform.py --image cover.png` (or `render.py --template audiogram`) puts the waveform over a still plate for an episode that has no picture, with `--platform` for the frame, `--title` and burnt-in captions. The image is a local file you give: nothing is fetched and no cover art is ever invented.\n- **Emoji in captions and titles (1.15).** `caption.py`\u002F`graphics.py --emoji-assets DIR` composites a PNG per emoji (Twemoji\u002FNoto naming, `1f389.png`) on top of the text, because drawtext cannot load a colour emoji font at all and an installed one does not prove libass will draw it in colour — `doctor --json`'s `fonts.emoji` answers that from a render probe. Without assets the run still succeeds and reports `mode: mono`. Nothing is ever downloaded.\n- **Indic and Thai text shaped correctly in titles and lower-thirds (1.15).** `graphics.py` renders Devanagari, Bengali, Tamil, Thai and Lao through libass automatically (`text_renderer: \"ass\"`), because drawtext never reorders matras or re-clusters marks; Arabic and Hebrew were already correct on a fribidi build. `--text-render drawtext` with such a script is refused, never rendered wrongly.\n- **Non-Latin text picks a font by script (1.12).** Japanese, Chinese, Korean, Arabic, Hebrew, Devanagari, Thai, Cyrillic and Greek cues, titles and overlays resolve a font file that covers them automatically, and a machine with no such font fails the job (`kind: input`) instead of rendering boxes. `doctor --json`'s `fonts.scripts` says which languages this machine can render; `--lang ja|ko` disambiguates Han-only text; an explicit `--font`\u002F`--font-file` is always kept.\n- **Silence detection finds nothing?** The default threshold is −35 dBFS. The tool prints a hint with the track's measured level; raise the threshold (`silence.py --threshold -25`) or shorten `--min-silence`.\n- **Sync results carry a confidence.** Below 0.3, or an offset near the edge of the analysis window, is probably wrong: enlarge `--analyze-seconds` or find a clap. Recordings over ten minutes from separate devices need `sync.py --fix-drift`.\n- **Outputs are never overwritten silently.** An existing output path is warned about today and refused from 2.0; set `FFMPEG_SKILL_NO_OVERWRITE=1` (the recommended agent setting) to get the refusal now and pass `--overwrite` where a replacement is intended.\n- **Long chains belong in a plan.** Three hand-chained re-encodes lose quality and are hard to change; `render.py` runs the whole edit from one JSON file, and `--dry-run` shows every ffmpeg command before anything is written.\n\n## FFmpeg compatibility\n\nThe tools need FFmpeg 5.0 or later and Python 3.9 or later (standard library only). What CI actually exercises on every pull request is FFmpeg 5.1.1 (static build), 6.1 (Ubuntu apt), 7.1 (Debian trixie apt), 8.x (macOS Homebrew) and 9.x (Windows gyan.dev), on Python 3.9 and 3.13 (the two ends of the supported range). The capability parser has been run against the listings of these builds:\n\n| FFmpeg | `-filters` row layout | Source |\n|---|---|---|\n| 5.1.1 | three flag characters, same as 6.x | johnvansickle.com static build on the Linux CI runner |\n| 6.1.1 | three flag characters: `..C acompressor A->A` | Ubuntu 24.04 apt, captured |\n| 7.1.x | same as 6.x | Debian trixie apt in a CI container (plus a constructed fixture in tests\u002F) |\n| 8.1.2 | two flag characters: `TS aap AA->A`, three-character legend, `------` separator | Homebrew on the macOS CI runner, captured |\n| 9.0.1 | same as 8.x, CRLF | gyan.dev build on the Windows CI runner, captured |\n\nFFmpeg 8 shortened the flag column of `ffmpeg -filters`. A parser anchored on the old width matches nothing on FFmpeg 8 and, if \"nothing matched\" is read as \"nothing installed\", reports every filter missing; that is what 0.9.0 did on macOS and Windows. Since 0.9.1 rows are recognised by their io-spec token (`A->A`, `AA->A`, `|->V`, `N->N`), so the flag width, the legend and the separator do not matter, and a listing that still cannot be read yields `unknown` rather than `missing`. The captured listings live in [tests\u002Ffixtures\u002F](tests\u002Ffixtures\u002FREADME.md) with their provenance; CI uploads each runner's listing and `doctor --json` as an artifact so a new layout is visible before it bites.\n\n## Tested on real footage\n\n**What is tested where.** The contract and the test suite (`tests\u002Ftest_contract.py`,\n`tests\u002Ftest_all.py`, which aggregates one module per tool group — `test_analysis.py`,\n`test_editing.py`, `test_audio.py`, `test_picture.py`, `test_delivery.py`,\n`test_orchestration.py` — over the shared footage in `tests\u002F_fixtures.py`) run on Linux, macOS\nand Windows on every pull request, minus the handful of POSIX-shim tests listed under\n[Development](#development). The real-device media corpus\n(`tests\u002Fcorpus.py`) has been run on Linux and macOS; the full corpus has **not** been run on\nWindows yet, and neither has an install by someone other than the maintainer been reproduced\nthere — [issue #143](https:\u002F\u002Fgithub.com\u002Fkajisho5\u002Fffmpeg-skill\u002Fissues\u002F143) tracks both. Treat the\nnumbers below as measured on Linux (and, where stated, macOS), not as a claim about every file\ntype on every OS.\n\n| Result | Measurement |\n|---|---|\n| **92 \u002F 92** | verification steps on a 10-file real-device corpus (GoPro, DJI, iPhone incl. Dolby Vision, Android screen recordings, HDR10, 24p, Tears of Steel), 0.8.0, local ffmpeg 6.1 |\n| **40 \u002F 40 within 10 ms** | `sync.py` offset detection, ±30 s offsets with gain, noise and EQ changes on real dialogue and music, 120 s windows (max error 1.1 ms); 60 s stress windows 95 % within 10 ms, 4 of 5 misses flagged by confidence |\n| **0 missed gaps** | `silence.py`, 20 cases with known gaps, ≤ 1 ms leftover silence |\n| **F1 0.97** | `scenes.py`, 53 hard cuts between single takes, precision 0.95, recall 1.00 at the default threshold |\n| **exact to the sample** | `cut.py --accurate` on WAV, FLAC (44.1 kHz) and AAC → WAV; WAV stream copy within 2 ms; AAC output +21 ms of encoder priming, reported as `codec_frame` (0.9.1) |\n| **72 \u002F 72** | agent runs of 24 prompts (12 English edits, 8 Japanese, 4 that must be declined), three repeats, graded by an independent model: routing, honest refusals and user's language 72\u002F72, report format 71\u002F72, visual check whenever the picture changed 24\u002F24 (0.8.4) |\n| **7 \u002F 8 routed** | eval 22 at 1.18.3 (2026-09-17, eight new symptom-only prompts for the five 1.18.0 flags, naming no flag, plus three repeats each of `cs1`\u002F`cs3`, Sonnet agent): after 1.18.1's SKILL.md routing rows and 1.18.2's matching README rows, `scenes.py --shots`, `cropdetect.py --motion-centre`, `silence.py --speech-aware`, `sync.py`'s N-source form and `multicam.py --switch energy` were all found from a symptom alone — including a Japanese and a Spanish variant — up from eval 21's 4\u002F9. The eighth prompt (`--filler` composed with `--speech-aware`) is an honest partial: the fixture has no real speech to transcribe, confirmed by installing `faster-whisper` mid-eval and re-testing by hand. `cs3` no longer rewrites the user's captions in any of 3 runs — the 1.17.3 fix holds through 1.18.1-1.18.3. Written up in `evals\u002Fresults\u002Fiteration-22.json` |\n| **8 \u002F 12 routed** | eval 21 at 1.18.0 (2026-09-17, one prompt per new analysis\u002Fmulticam flag plus a `cs1`\u002F`cs3` recheck, Sonnet agent): every result was correct where the agent found the right script — both sync offsets, the shot label, the multicam switch point and its render-project mapping, both 1.17.2\u002F1.17.3 rechecks — but 4 of 12 prompts hit a script the agent could not find, because SKILL.md named none of the five new flags (confirmed by grep). Three refused honestly rather than fabricate a number; one reached a correct answer without the intended flag, by luck of one fixture's silence durations. 1.18.1 (2026-09-17) is the SKILL.md fix, since evaluated by eval 22. Written up in `evals\u002Fresults\u002Fiteration-21.json` |\n| **20 \u002F 20** | 1.17.2 run (2026-09-14, the eight caption prompts of eval 19, six of them three times, Sonnet agent, regex grader + an Opus grader that opened every contact sheet and counted the lines per cue): routing 19\u002F19 act runs, honest 18\u002F20 with 0 false successes and 0 raw ffmpeg calls, report format 20\u002F20, user's language 20\u002F20, trigger set 49\u002F50 (one judge flip on a file-less prompt), Opus quality mean 4.25 (3.65 at eval 19). The picture is fixed: 0\u002F20 runs stack one word per line against 12\u002F12 template runs at eval 19 on the same cues; the Style row at TikTok geometry is now `…,54,151,420,1`, the fitter's `size_used` (15 on TikTok, 16 on Shorts, the 13 floor for the Spanish cues) is what the frame shows, and report and sheet agree in 18\u002F20 runs. What is left is not the typesetter: `cs3` rewrote the user's captions for the fourth iteration running, one `cs1` run raised `max_lines` to 4 to avoid a shrink and drew four-line stacks, and `cs2`'s 32-letter Spanish word still leaves the frame at the size floor (disclosed 3\u002F3). Written up in `evals\u002Fresults\u002Fiteration-20.json` |\n| **26 \u002F 26** | 1.17.1 run (2026-09-14, targeted re-run of the 18 prompts eval 18's follow-up named, plus three repeats each of the four caption-size prompts, Sonnet agent, regex grader + a full Opus grader over all 26 runs, every PNG opened and every written output re-probed): routing 23\u002F26, honest refusals and failures 24\u002F26 with 0 false successes and 0 raw ffmpeg calls, report format 26\u002F26 with the third label gone, user's language 26\u002F26, trigger set 50\u002F50, Opus quality mean 3.65. 1.17.1's fix holds — `--fit-size` now fires on the `render.py --template` path in 12\u002F12 caption runs (24 → 16, `dl4` to the 13-unit floor, `split` 0, `text_unchanged` true, identical across repeats), the beat, filler and `--jobs` prompts route on the first try, and `bt2` quotes its measured 0.184 confidence instead of denying the capability exists. The honest part: the picture is unchanged. `caption.py`'s `write_ass` writes the platform's *vertical* safe margin into `MarginL`, `MarginR` and `MarginV` alike (tiktok 63 ASS units → 420 px), so at `PlayResX` 1080 the text column is 240 px and libass wraps every word — the fitter budgets `play_w × 0.9`, which is why `split: 0` is true of the ASS text and false of the frame. It is the `--animate`\u002F`--karaoke` path only, present since 1.14, and it explains eval 17's and eval 18's \"one word per line\" too; 1.17.2 is the patch and the finding is written up in `evals\u002Fresults\u002Fiteration-19.json` |\n| **100 \u002F 100** | 1.17.0 run (2026-09-14, one pass per prompt, Sonnet agent, regex grader + focused Opus grader on 28 runs, every written output re-probed, `check.py` re-run on every delivery output) on the set grown to 100 prompts (caption size fitting, beat-synced cuts, filler removal, batch `--jobs`, render `--cache`): routing 95% over the 64 act prompts, honest refusals and failures 22\u002F25 with 0 false successes and 0 raw ffmpeg calls, report format 98\u002F100 (two runs label an honest partial result with a third label), user's language 100\u002F100 across seventeen languages, visual check 24\u002F24, real execution 6\u002F6 with honest failure 5\u002F5, trigger set 50\u002F50 including all five new 1.17 prompts, Opus quality mean 3.71. The honest part: `--fit-size` is unreachable on the template path (`render.py` forwards the platform table's caption size as an explicit `--size`, so the fitter declines to shrink a size it thinks the user chose, and the project schema rejects `fit_size` outright — only the one run that called `caption.py` by hand got 24 → 16, `split` 0), and SKILL.md names none of the 1.17 features, so beats, filler and `--cache` were each used in one run at most — 1.17.1 is the patch and the finding is written up in `evals\u002Fresults\u002Fiteration-18.json` |\n| **90 \u002F 90** | 1.16.0 run (2026-09-14, one pass per prompt, Sonnet agent, regex grader + focused Opus grader on 30 runs, chapters and subtitle streams re-probed, check.py re-run on every delivery output) on the set grown to 90 prompts (audiogram, auto chapters, multi-language tracks, caption breaking): routing 90\u002F90, honest refusals and failures 90\u002F90 with 0 false successes and 0 raw ffmpeg calls, report format 89\u002F90 (one `Done (partially):`), user's language 90\u002F90 by regex (89\u002F90 by Opus), audiogram 2\u002F2 with the cover behind the waveform and nothing fetched, auto chapters 2\u002F2 with `Chapter N` titles only, delivery 16\u002F16 platform pass, trigger set 45\u002F45, Opus quality mean 4.17. The honest part: the phrase breaker never gets to act at the platform caption sizes (a five-word cue does not fit two lines at TikTok size, so the split is byte-identical to 1.15.1), Thai still breaks inside words, and a katakana word was split — 1.16.1 is the patch and the finding is written up in `evals\u002Fresults\u002Fiteration-17.json` |\n| **82 \u002F 82** | 1.15.0 run (2026-09-13, one pass per prompt, Sonnet agent, regex grader + focused Opus grader on 28 runs, stills extracted inside the emoji window, check.py re-run on every delivery output) on the set grown to 82 prompts (emoji captions and title cards, a Hindi and a Thai lower-third): routing 82\u002F82, honest refusals and failures 82\u002F82 with 0 false successes and 0 raw ffmpeg calls, report format 82\u002F82 (both iteration-15 label defects closed: `dl8` and `he2` now carry one `Failed:`), user's language 82\u002F82 by regex (81\u002F82 by Opus: one Spanish report with three English labels), non-Latin glyphs 11\u002F11 (Devanagari through `graphics.py` is fixed; Thai lower-third and captions correct), emoji visible in colour in 3\u002F3 runs given PNG assets and reported monochrome in the one that was not, visual check 23\u002F24, delivery 12\u002F13 one encode and 13\u002F13 platform pass, trigger set 40\u002F40, Opus quality mean 4.68. Still open: the caption breaker splits phrases (`dl1`, `dl4` unchanged) — queued for 1.16.0. Tokens per run flat at 73.3k on the same 76. Details in `evals\u002Fresults\u002Fiteration-16.json` |\n| **76 \u002F 76** | 1.14.0 run (2026-09-13, one pass per prompt, Sonnet agent, regex grader + focused Opus grader on 26 runs, check.py re-run on every delivery output): routing 76\u002F76, honest refusals and failures 76\u002F76 with 0 false successes and 0 raw ffmpeg calls, report format 76\u002F76, user's language 76\u002F76 by regex (75\u002F76 by Opus: one Spanish report with three English labels), visual check 18\u002F18, trigger set 38\u002F38, Opus quality mean 4.58. The delivery templates did their job: 12 of 13 delivery requests went through `render.py --template`, finished in one encode (was 3 of 7) and all 13 pass their platform check (was 7 of 8). Tokens per run flat at 73.4k. Details in `evals\u002Fresults\u002Fiteration-15.json` |\n| **76 \u002F 76** | 1.13.0 run (2026-09-13, one pass per prompt, Sonnet agent, regex grader + focused Opus grader) on the set grown to 76 prompts: 18 in Thai, Hindi, Hebrew, Russian, Greek, Vietnamese, Indonesian, Turkish and Italian, and 8 delivery requests (TikTok, Reels, Shorts, LinkedIn, Douyin, podcast): routing 76\u002F76, honest refusals and failures 76\u002F76 with 0 false successes and 0 raw ffmpeg calls, report format 76\u002F76, user's language 76\u002F76 across seventeen languages, visual check 18\u002F18, trigger set 38\u002F38, Opus quality mean 4.65 over the 26 new runs. One real defect found: Hindi through `graphics.py` (drawtext) comes out wrong-shaped even though the font covers Devanagari; captions through libass are fine (queued for 1.15.0). Four delivery runs spent a second encode for loudness, which 1.14.0's templates address. Tokens per run flat at 72.3k. Details in `evals\u002Fresults\u002Fiteration-14.json` |\n| **50 \u002F 50** | 1.12.0 run (2026-09-13, one pass per prompt, Sonnet agent, regex grader + focused Opus grader) on the set grown to 50 prompts with two each in Chinese, Korean, Spanish, Portuguese, French, German and Arabic: routing 50\u002F50, honest refusals and failures 50\u002F50 with 0 false successes and 0 raw ffmpeg calls, report format 50\u002F50, user's language 50\u002F50 across nine languages, visual check 13\u002F13, trigger set 29\u002F29, Opus quality mean 4.83. Every non-Latin caption and lower-third picked a covering font by itself and rendered real glyphs (Arabic shaped and right-to-left); tokens per run unchanged at 72.3k. Details in `evals\u002Fresults\u002Fiteration-13.json` |\n| **36 \u002F 36** | 1.11.1 re-run (2026-09-13, one pass per prompt, Sonnet agent, regex grader + focused Opus grader): routing 36\u002F36, honest refusals and failures 36\u002F36 with 0 false successes and 0 raw ffmpeg calls, report format 36\u002F36, user's language 36\u002F36, visual check 8\u002F8, trigger set 22\u002F22, Opus quality mean 4.75. The 1.11.1 wording did what it said (`doctor` before a job 23 of 36 runs → 0, `--json-brief` 4 → 23) and tokens per run stayed flat at 71.8k, because about 64k of every run is the host's own context; the token-diet theme closes here. Details in `evals\u002Fresults\u002Fiteration-12.json` |\n| **36 \u002F 36** | 1.11.0 re-run (2026-09-13, one pass per prompt, Sonnet agent, regex grader + focused Opus grader): routing 36\u002F36, honest refusals and failures 36\u002F36 with 0 false successes and 0 raw ffmpeg calls, report format 36\u002F36, user's language 36\u002F36, visual check 8\u002F8, trigger set 22\u002F22. First iteration to measure tokens per run: mean 72.2k against 68.7k at 1.10.0, mostly a fixed per-run floor the skill does not control (a refusal run that only reads SKILL.md costs about 64k), plus `doctor` on 23 of 36 runs; 1.11.1 rewords step 0 and iteration 12 re-measures. Details in `evals\u002Fresults\u002Fiteration-11.json` |\n| **108 \u002F 108** | 1.10.0 re-run (2026-09-13, three passes per prompt, Sonnet agent, regex grader + focused Opus grader): routing 108\u002F108, honest refusals and failures 108\u002F108 with 0 false successes and 0 raw ffmpeg calls, report format 108\u002F108 by both graders (the harness now names the five labels), user's language 105\u002F108 (every Japanese request in Japanese; 3 English requests drifted to Spanish or Portuguese), visual check 21\u002F24, trigger set 22\u002F22; real-device corpus 101\u002F101 steps PASS. Details in `evals\u002Fresults\u002Fiteration-10.json` |\n| **108 \u002F 108** | 1.9.0 re-run (2026-09-13, three passes per prompt, Sonnet agent, regex grader + independent Opus grader): routing 108\u002F108, honest refusals and failures 108\u002F108 with 0 false successes and 0 raw ffmpeg calls, visual check 23\u002F26, user's language 105\u002F108 (every Japanese request answered in Japanese; 3 English requests drifted to Spanish), report format 108\u002F108 by regex (65\u002F108 by the stricter grader, which now counts any missing label), trigger set 22\u002F22; 12 of 15 platform jobs were one encode and `render.py` rendered once in 3\u002F3 (was 1\u002F3). The 1.9.0 time grammar was not used by any agent. Details in `evals\u002Fresults\u002Fiteration-9.json` |\n| **108 \u002F 108** | 1.8.0 re-run (2026-09-12, three passes per prompt, Sonnet agent, regex grader + independent Opus grader that re-probed 22 outputs): routing 108\u002F108, honest refusals and failures 108\u002F108 with 0 false successes and 0 raw ffmpeg calls, visual check 22\u002F24, report format 108\u002F108 by regex (91\u002F108 by the stricter grader: 'What\u002Fhow:' in place of Steps:), user's language 98\u002F108 by the stricter grader (Japanese labels-only reports counted), trigger set 22\u002F22; 13 of 14 platform exports used `--normalize` and platform jobs went from three encodes to one; r04\u002Ff01 now answered in the request's language 5\u002F6 (was 0\u002F6). Details in `evals\u002Fresults\u002Fiteration-8.json` |\n| **108 \u002F 108** | 1.7.0 re-run (2026-09-12, three passes per prompt, Sonnet agent, regex grader + independent Opus grader that re-probed 24 outputs): routing 108\u002F108, honest refusals and failures 108\u002F108 with 0 false successes and 0 raw ffmpeg calls, visual check 25\u002F25, report format 108\u002F108 by regex (104\u002F108 by the stricter grader), user's language 101\u002F108 (six English refusals answered in Spanish or Portuguese, one Japanese request in English); trigger set 22\u002F22. Iteration-6 fixes held (fade-in only, Japanese audio trims, music no longer shortens the video). Details in `evals\u002Fresults\u002Fiteration-7.json` |\n| **36 \u002F 36** | 1.4.15 re-run (2026-09-12, one pass per prompt, Sonnet agent, regex grader + manual review): 24-prompt set routing 20\u002F20, honest refusals 5\u002F5, visual check 8\u002F8, report format 25\u002F25, user's language 9\u002F9; exec set real execution 6\u002F6, honest failure on bad inputs 5\u002F5 with 0 false successes, audio-as-audio 3\u002F3, one Japanese report with English labels; trigger set 22\u002F22. Both iteration-5 defects gone (no raw ffmpeg fallback, music no longer shortens the video). Details in `evals\u002Fresults\u002Fiteration-6.json` |\n| **36 \u002F 36** | 1.4.0 re-run (2026-09-11, one pass per prompt, Sonnet agent, regex grader + manual review): 24-prompt set routing 20\u002F20, honest refusals 5\u002F5, visual check 8\u002F8, user's language 8\u002F9; exec set real execution 6\u002F6, honest failure on bad inputs 5\u002F5 with 0 false successes, audio-as-audio 3\u002F3; trigger set 22\u002F22. Details and the six findings in `evals\u002Fresults\u002Fiteration-5.json` |\n| **6 \u002F 6** | 0.9.1 audio evals (audio join, extraction, track selection, sample-accurate trim, typed dynamics; 2 in Japanese): routing, report format and audio-as-audio handling 6\u002F6 |\n\n```bash\npython3 tests\u002Fcorpus.py --fetch --verify     # ~1.4 GB download, then verify (slow on 4K)\npython3 tests\u002Fbench_sync.py --cases 100\npython3 tests\u002Fbench_silence.py\npython3 tests\u002Fbench_scenes.py\n```\n\nBenchmarks live in `tests\u002Fbench_*.py`, agent evals in [evals\u002F](evals\u002F), results by iteration in `evals\u002Fresults\u002F`.\n\n## Install\n\n```bash\nnpx ffmpeg-skill              # Claude Code   → ~\u002F.claude\u002Fskills\u002Fffmpeg-skill\nnpx ffmpeg-skill --cursor     # Cursor        → ~\u002F.cursor\u002Fskills\u002Fffmpeg-skill\nnpx ffmpeg-skill --codex      # Codex         → ~\u002F.agents\u002Fskills\u002Fffmpeg-skill (Cursor reads this location too)\nnpx ffmpeg-skill --all        # all three\nnpx ffmpeg-skill --project    # this project  → .\u002F.claude\u002Fskills\u002Fffmpeg-skill\nnpx ffmpeg-skill --dir .\u002Fmy-skills\nnpx","ffmpeg-skill 是一个专为 AI 编程代理（如 Claude Code、Cursor）设计的本地化音视频处理工具包，通过封装 FFmpeg 命令提供零依赖、无云服务、无需 API 密钥的媒体操作能力。核心功能包括智能字幕生成与动画（含卡拉OK高亮）、自适应画面裁剪与重构帧（如横屏转竖屏）、静音段自动检测与移除、响度标准化等，全部基于 Python 标准库调用本地 FFmpeg 进程。适用于需要在本地环境中快速集成音视频预处理\u002F后处理能力的 AI 代理工作流、自动化内容生成及边缘侧媒体编辑场景。",2,"2026-09-06 02:30:03","CREATED_QUERY"]