Open source · Claude Code

inkwell

A product URL in, ad videos in every aspect ratio out. Or a narrated launch film.

It reads a product page for the imagery and brand colours that page already publishes, then renders the same ad at every size a platform accepts. Each ratio is laid out separately rather than letterboxed from one master, because a vertical ad with bars down the top and bottom throws away the part of the frame that does the work. For a launch, it makes a narrated film instead: one voice, a scored take per line, and sound design under every beat.

git clone https://github.com/debashisn94/inkwell.git

MIT licensed. Needs Node and ffmpeg, plus Python and VoxCPM for narration. No API keys, no accounts.

One definition, four cuts.

Type scales off the short edge, so a headline is not enormous at 16:9 and unreadable at 9:16. Safe areas match what the platform actually covers. The stack direction flips between portrait and landscape, so neither is a crop of the other.

  • Reel 9:16 1080 x 1920 Reels, Shorts, TikTok, Stories
  • Feed 4:5 1080 x 1350 Instagram and Facebook feed
  • Square 1:1 1080 x 1080 Feed, LinkedIn
  • Wide 16:9 1920 x 1080 YouTube, LinkedIn, X, pre-roll
Vertical and wide, rendered from the same file and playing in sync.

Three commands, and the one that matters is the middle one.

  1. 01

    Research the product

    node tools/new-ad.mjs https://your-product.com --out=ad

    Reads the page and writes brand.json: name, positioning, headings, the imagery it publishes, and a palette sampled from that imagery. Then drafts an ad.json with every line marked for rewriting.

  2. 02

    Write the ad

    edit ad/ad.json

    Four scene types: hook, feature, proof, call to action. This is the part that decides whether the ad works, and it is the part the tool deliberately does not finish for you.

  3. 03

    Render every size

    cd ad && npm install && node ../tools/render-ad.mjs

    Bundles once and renders all four formats. Adding a fifth is one entry in a list, because the layout engine derives type scale, safe areas and stack direction from the dimensions.

New in v0.3

Or a narrated launch film.

An ad has fifteen seconds. A launch needs two minutes: the problem shown happening, why the obvious fix fails, the reveal, how it works, and the offer. Same research step, a different output, voiced and sound-designed.

Holt Teams, 1:51, sound on. Every frame and every sound comes from one film.json in the repo.
  • 12scene types, from a live terminal to a lock slam and an architecture flow
  • 1narrator for every line, so the voice never changes mid-film
  • 15synthesized sounds, each scene firing its own on its own beats
  1. 01

    Write the film

    node tools/new-film.mjs https://your-product.com --out=film

    Scaffolds the project, pulls the palette from the product’s own site, and lays out a structure: problem, the obvious fix and why it fails, reveal, how it works, offer, action. Each scene carries its on-screen copy and one line of narration.

  2. 02

    Voice it

    python ../tools/film-voice.py

    Designs one narrator, then clones every line from it with its own delivery note. Each line gets two or more takes, each take is transcribed back and checked against the narrator’s voice, and the most expressive take that passes is kept.

  3. 03

    Render picture and sound

    node ../tools/render-film.mjs --stems

    One pass renders the picture, the narration and the sound design, then masters it to -16 LUFS. The stems come out separately, so music goes in under the voice without re-rendering anything.

Scene types, the voice scoring, and mixing music under the stems are in the film guide.

Six decisions that took more than one attempt.

Each of these was wrong first, in a way that looked plausible until it was measured.

It reads the page head, not the images

Most product sites render their imagery in JavaScript or as CSS backgrounds, so scraping image tags returns nothing on exactly the polished sites worth advertising. The og:image is reliably present, correctly sized, and chosen by the owner to represent the product.

The palette comes only from the brand’s own assets

Sampling every image on a page sounds more thorough and is wrong. On a portfolio or a customer-logo strip those are other companies’ marks. Tested against a site carrying client logos, the extracted accent came back as the client’s teal instead of the brand’s amber.

Colour is read at high resolution or not at all

Sampling a 1200 by 630 banner on a 32 by 32 grid averages each cell across roughly 37 pixels of source, which is wider than the lettering in most logos. The accent is destroyed before any of the colour logic runs. It samples at 260 and the extracted colour lands within a couple of percent of the real one.

One narrator, not eleven

Designing each line from a written voice description sounds like the natural way to direct a narrator. Tested on the first film, every line came back as a different person. The narrator is now designed once and saved, and every line is cloned from that recording. Delivery changes per line. The voice does not.

The narration sets the edit

Scenes were first timed by hand, and every re-voiced line broke the timing of everything after it. Now each scene lasts as long as its animation needs or its line needs plus a breath, whichever is longer. Re-voice a line and the film retimes itself.

Assets keep their own shape

An og:image is a 1.91:1 banner, usually with the product name set across it. Fitting that into a portrait box with a cover crop takes the wordmark off both ends, so every asset is measured and given a box that matches it.

What it does not do.

  • It drafts ad copy from the page and marks every line for rewriting. It can tell you what a product calls itself. It cannot tell you who it beats or what the objection is, and those are what make an ad work.
  • Every ad shares one visual grammar: text on a gradient with a card holding the product’s own image. Well executed, and recognisable. Forty of these would look like forty of the same thing.
  • There is no device. A UI assembling itself, a before and after wipe, a real screen recording. Those have to be designed per product rather than derived from a URL.
  • Films are 16:9 only. Their diagrams and side-by-side layouts need a wide canvas, and a vertical cut of a film is a different edit, not a reflow.
  • Narration runs locally on VoxCPM, which needs a capable GPU or Apple silicon and Python. The takes are scored automatically, but the score cannot hear a flat read. Someone still has to listen.
  • Remotion is free for individuals and teams of three or fewer, and needs a company licence above that.

There is a write-up of the part that took three goes, including why the first version measured perfectly and still looked like a slide deck: read the build notes.

Point it at a product you already know, and see whether the ad it drafts is worth rewriting.