Short answer. A second of generated video costs five cents on the cheap models and fifty on the expensive ones. An image runs three or four cents. An extra minute of voice is under a dollar. All of these prices are published. The conclusion people draw from them is wrong: one operation out of ten got cheaper, and it was never the expensive one. The number to count is the cost of a published unit, and generation is a couple of per cent of it.
Generation itself now costs cents. Everything around it costs exactly what it always did.
What does generation actually cost?
Cents. Everyone selling access on demand publishes their prices openly, and an afternoon is enough to put them into a single table.
| Item | Cheap | Expensive | What doubles it |
|---|---|---|---|
| A second of video | 5 cents, Wan 2.5 | 50 cents, Sora 2 Pro at 1080p | audio and resolution |
| An image | 2 or 3 cents, Qwen and Seedream V4 | 4 cents, Flux Kontext Pro | size in megapixels |
| A minute of voice | 17 cents | 36 cents | your plan, not the quality |
Prices come from the fal.ai, Replicate and ElevenLabs price lists as of September 2026. They move, so open them again before you plan a budget.
The spread has a simple explanation. Cheap models generate a short clip without audio at middling resolution, which is enough for a feed. Expensive ones offer a long take, audio and voice control, which is what you need when somebody watches the whole thing. Kling 2.6 Pro shows it in one line: seven cents a second without audio, fourteen with it, and 16.8 with voice control on top. Audio exactly doubles the bill.
In plain terms: a thirty-second 1080p clip with audio costs twelve dollars on Veo 3.1, four dollars and twenty cents on Kling 2.6 Pro, and a dollar fifty on Wan 3.0 at 480p. Twenty images for a month of posting cost less than a dollar. Narrating a half-hour course comes to about five dollars fifty.
Why is that the wrong invoice to count?
Because several discarded units sit behind every published one, and both figures stay small anyway.
Let us add the attempts in. Getting one usable clip normally takes five goes, sometimes ten: wrong framing, wrong emotion, bent hands, text gone crooked. Call it eight. A thirty-second clip on Kling 2.6 Pro with audio costs four dollars and twenty cents, so eight attempts come to thirty-four dollars.
Now for the comparison this was all leading to. Thirty-four dollars is roughly one hour from the person who wrote the brief, reviewed the options and assembled the final cut. One hour. A clip does not take an hour, though. It usually takes a working day, and several other people are involved in that day: whoever approves it, whoever fixes the copy, whoever puts it in the publishing plan.
The consequence is worth accepting before you pick a model. Saving on generation costs is pointless. Across a month the difference between a cheap model and an expensive one comes to tens of dollars, while one extra round of approval comes to hundreds.
So what is expensive?
Everything around the generation. The brief, the selection, the approval, the redo.
The brief. Until somebody writes down what is needed and why, generation produces something pretty that misses the mark. This is the one step you cannot skip, and it is entirely human.
The selection. Somebody picks one of the eight, and that is a person too. What lands with your customer is not something the model knows: it has seen neither your sales nor your rejections. You can give it that knowledge, by collecting the history of those decisions and passing it in with the task. Almost nobody does, because that means an extra configuration layer, but the option exists and with it the selection stops being entirely manual.
The approval. The most expensive step and usually the least noticed. Two rounds of comments from three people cost more than a whole month of generation, because they are counted in other people's working hours.
The redo. It happens where the brief was written badly, which brings us back to the first item. In our experience that is the most common reason for a second and third round: the model did not get it wrong, the target was set so that hitting it was impossible.
Three of the four can be put into a system. A brief becomes a template, selection becomes stated criteria, approval becomes one round with a set deadline. Nothing that requires taste fits into one, and replacing it is not the point. That is how our content factory is built, and the machine is the cheapest part of it.
How do you budget?
Cost per published unit. Everything spent in a month, divided by everything that went out.
Four things go into the spend: subscriptions and generation, your own people's hours, contractor hours, and the cost of delay. That last one usually goes uncounted and it is real. A post that went out three weeks after the moment that prompted it cost full price and returned nothing. Delay appears on no invoice, which is why nobody sees it until somebody starts writing the date of the trigger next to the date of publication.
In the first month, two numbers are enough. How many units went out and how many hours went into them. A month later you have a cost per unit, and it is almost always several times higher than the one budgeted from a tool's price list.
One sign tells you the system paid off. Cost per unit falls while the number of units rises. If only the number rises, you have hired more hands. If only the cost falls, you are probably publishing worse work. Read the two together, because either one alone is easy to game.
The honest conclusion. The real saving here is not in the price of a second. It is in writing the brief once instead of rewriting it every month, and in keeping approval to one round. A five-cent machine decides none of that, and cannot. It performs exactly the operation that took the least time to begin with.
Danil Ivanov
Founder, KAIVIX
Builds AI systems for companies in the UAE and beyond.