I Rendered a Month of Video in One Night on Machines I Already Owned

Thirty-nine short videos and four long ones, voiced and rendered in a single night on hardware I already had. The bottleneck was the one machine that could run the voice model, plus two failures that reported success.

I Rendered a Month of Video in One Night on Machines I Already Owned · Products Decoded

In one night, 39 short videos and 4 long-form YouTube videos went from scripts to finished files with captions, sitting in a delivery folder. Nothing new was bought to make that happen. Three machines I already had did the work, and every file came out the other end of an automated frame check.

The interesting part is not the count. It is that I planned the whole night around the wrong resource, and the two things that nearly ruined it both reported success while failing.

I assumed the expensive part would be rendering

Video looks like a compute problem. My mental model going in was that the render queue would be the wall, and the job was to throw as many cores at it as possible.

The three machines: a 16GB M1 Pro laptop, an older 8GB M1 that stays on as a worker, and a rented Linux box. Straightforward split, I thought. Laptop coordinates, the other two grind.

Then I looked at the voice step. The voice model runs locally on Apple Silicon, it clones a voice, and it costs nothing per clip because it never leaves the machine. It also only runs on one of the three boxes, and it runs one clip at a time.

That single sentence set the entire schedule. Voice was the critical path, all night, and every other decision had to arrange itself around keeping that one process fed.

The rented box cannot talk at all

Naturally I tried to move voice onto the rented box first, because that is the machine with headroom and the one whose job is to be busy while I sleep.

It does not work there. The model upcasts to float32 on that hardware, memory climbs, and it dies at 9GB. Eight minutes in, zero output. Not slow, nothing at all. I tried it, it failed the same way twice, and I stopped. Some things are worth a third attempt and this was not one of them, though I will admit I am not fully sure whether a smaller precision setting would have rescued it and I did not have the night to find out.

So the laptop does voice, serially, for as long as it takes. The laptop also renders nothing, ever, because anything that steals cycles from voice steals them from the whole night.

The little machine was the faster renderer

Here is the part that would have cost me hours if I had trusted my assumptions. The 8GB M1, the small one that mostly sits there running a bot, renders faster than the rented Linux box.

I had it the other way around in my head. If I had assigned the queue by intuition, the fastest renderer in the house would have been sitting at the back doing a third of the work. I found out by measuring one job on each and comparing, which took about the time it takes to make tea.

The scheduling shape that fell out of all this: voice runs continuously on the laptop, one clip after another. Each finished audio file drops a render job into a queue. The two renderers pull from that queue in parallel and neither of them ever waits for the other, only for voice. Voice finishes last on the last clip, and the night ends whenever it says so.

Once the critical path is clear, everything else in the pipeline gets cheap. Captions, the frame check, moving finished files into a delivery folder, all of it runs on machines that would otherwise be waiting, so none of it shows up on the clock. The whole night collapses into one question, which is whether the voice queue ever went idle. Time lost there is time the night does not get back, and time lost anywhere else mostly does not matter.

Half the render jobs did nothing and reported nothing

First silent failure, and this one is a classic that I still walked into.

node was not on the small M1’s non-interactive SSH PATH. When you log into a machine and type a command, your shell reads your profile and knows where everything lives. When a script pushes a command over SSH, that is a non-interactive shell, and it does not necessarily read the same files. The binary that exists perfectly well when you are sitting in the terminal is simply not there when a script asks for it.

So roughly half the render jobs went out, found no interpreter, and no-opped. Quietly. The queue kept draining and the pipeline kept moving, because from the coordinator’s point of view the job had been dispatched and had come back.

A smoke test caught it. One job, one machine, then go and look at the file it was supposed to produce with your own eyes before starting the batch. That is the entire practice and it is boring and I do it now because that night taught me what the alternative costs.

A crash would have been kinder. A crash produces a stack trace and a red line and someone goes and reads it. Silence produces a delivery folder that is missing files nobody has counted yet.

My QA tool passed everything because a flag hid the evidence

Second silent failure, and this one I am fonder of because it is genuinely funny.

I had written a small checker to sample frames out of each finished video and confirm nothing was blank or broken. It ran across the batch and passed everything. Which is exactly what you want to see, and exactly what I wanted to see at that hour.

The tool parsed the output of an ffmpeg command. That command was being run with -v error, which suppresses everything below error level, including the exact output the tool was reading. No output came back. No output parsed as no problems found. Every file passed for the same reason a smoke detector with the battery out never goes off.

The check I should have run before trusting any of it: hand the QA tool a file you know is broken and confirm that it fails. If your verifier has never once said no, you do not actually know that it can.

The pattern behind both of these bugs is the same, and it has very little to do with SSH or with ffmpeg. A step in a pipeline can fail in a way that produces no output at all, and an orchestrator that only checks whether the step returned will happily read that as success. Every automated pipeline I have built has eventually grown a bug of this shape. The fix is never cleverer error handling. Go and look at the artifact with your own eyes. Did the file appear, is it roughly the size you expected, does it contain the thing you asked for.

They did all pass the frame check in the end, after the flag came off and the thing could actually see what it was grading.

The kernel had to be told what to kill

The rented box is not a scratch machine. It runs live products that other people use, in containers, on the same host I was about to hammer with video encoding all night.

Every render job on that box ran as a systemd unit with a memory cap and OOMScoreAdjust=1000. When Linux runs out of memory it picks a process to kill, using a score. That setting pushes the render jobs to the top of the list, so when memory gets tight the kernel reaches for the video job first and leaves the production containers alone.

This got proven the same night. A job breached its cap, the kernel killed it, and the thing that died was the render. Nothing user-facing moved. I had a failed job to re-queue in the morning, which is the correct outcome and the cheapest possible version of that lesson.

If you are borrowing production hardware for a batch job, do this before the batch, not after the first incident. The default behaviour is the kernel choosing on your behalf, and the kernel does not know which of your processes has customers attached to it.

Do not edit the script after you chunk it

One more rule from that night, learned the hard way and cheap to pass on.

The voice step chunks a script into beats before it starts speaking. Once that has happened, going back and editing the script re-indexes everything downstream of the edit. A one-word fix in the middle is not a one-word fix. It shifts every beat after it, and the audio and the timing behind them stop lining up with what you think they are.

So the script gets frozen before chunking. If it is wrong, it gets fixed before, or it ships with the wrong word and gets fixed in the next batch.

The cost of the night was zero on top of what I already pay. Voice is local, the laptop was already mine, the rented box was already rented for other reasons, and the small M1 has been running in a corner for months. What I actually spent was attention, in two places: a smoke test before the batch, and a deliberately broken file fed to my own QA tool.

The ceiling is still voice. One machine, one clip at a time, and a night is only as long as a night. The obvious answer is a second Apple Silicon box to double the voice throughput, and I keep not buying one, partly because I am not convinced a month of video should be produced in a night at all. That was a fun thing to prove and a strange thing to make routine.