What live captioning actually costs in 2026

August 27, 2026

Every automatic speech recognition API bills you for silence.

Not as a trick. It is just how metered transcription works. You open a streaming session, the meter runs, and it keeps running through the dead air between plays, through the ad break, through the 3am infomercial nobody is watching. A linear channel runs 8,760 hours a year. Your invoice reflects 8,760 hours, whether the audio contained speech or a test tone.

For one channel that is fine. It is a rounding error against transmission costs. The problem starts when someone asks what it would cost to caption everything: every channel, every language, every regional feed. At that point the per-minute rate on the pricing page stops being a rate and becomes a multiplier, and finance wants to know why the AI line item grows every time programming adds a service.

This post works through the actual numbers. Cloud API list pricing as it stands in August 2026, what self-hosting genuinely costs once you include power and rack space and the engineer who maintains it, and where the crossover sits. All arithmetic is shown so you can substitute your own figures.

What does cloud ASR cost per minute in 2026?

Raw transcription has gotten cheap. Here is the published list pricing, retrieved 17 to 18 August 2026. Streaming and batch are priced separately almost everywhere, and streaming carries a premium of roughly 1.5 to 2 times.

Provider

Batch

Streaming / real-time

AssemblyAI Universal-2 / Universal Streaming

$0.15/hr

$0.15/hr

AssemblyAI Universal-3.5 Pro

$0.21/hr

$0.45/hr

Deepgram Nova-3

$0.0043/min ($0.26/hr)

$0.0077/min ($0.46/hr)

Speechmatics Pro (from)

$0.129/hr

$0.402/hr

ElevenLabs Scribe v2

$0.22/hr

$0.39/hr

Mistral Voxtral Transcribe 2

$0.003/min ($0.18/hr)

sub-200ms tier

OpenAI gpt-4o-transcribe

$0.006/min ($0.36/hr)

~$0.017/min ($1.02/hr)

OpenAI gpt-4o-mini-transcribe

$0.003/min ($0.18/hr)

not offered

Google Cloud STT v2 (Chirp)

$0.016/min, ~$0.003/min dynamic batch

$0.016/min ($0.96/hr)

Microsoft Azure AI Speech

$0.18/hr

$1.00/hr

AWS Transcribe

$0.024/min tier 1, falling to $0.0078/min above 5M min/mo

same tiers

Four things in the fine print matter more than the headline rate.

Azure charges continuous language identification as an add-on. Real-time standard is $1.00 per hour, and enhanced features including diarization and continuous language ID add $0.30 per hour. If you are monitoring a multi-language feed and need to know which language is being spoken, your real rate is $1.30, not $1.00.

AssemblyAI closes streaming sessions after three hours and bills the full session. If your process does not terminate cleanly you pay for three hours regardless of how much audio you sent. On a 24/7 feed that is eight reconnections a day, each one a billing event.

AWS bills a 15-second minimum per request and resets its volume tiers monthly. A channel that reaches tier 2 pricing in week three drops back to tier 1 on the first of the month.

Google's Dynamic Batch tier is roughly 75% cheaper than standard batch in exchange for up to 24-hour turnaround. For archive work that is free money. For anything live it is irrelevant.

One caveat on AWS specifically. Its published structure has been tiered per-minute for years, but at least one pricing tracker now reports flat batch and streaming rates with a 60-second minimum. Check the live AWS page before you build a model on it.

Why does per-minute pricing get expensive so fast?

Because a linear channel runs 8,760 hours a year and the arithmetic is just rate multiplied by hours.

One channel, streaming raw ASR, list pricing, no negotiated discount:

Provider

$/hr

Per channel-year

AssemblyAI Universal Streaming

$0.150

$1,314

Speechmatics real-time (from)

$0.402

$3,522

AssemblyAI U-3.5 Pro Realtime

$0.450

$3,942

Deepgram Nova-3 streaming

$0.462

$4,047

Google STT v2 real-time

$0.960

$8,410

Azure real-time standard

$1.000

$8,760

OpenAI realtime

$1.020

$8,935

Azure real-time plus continuous language ID

$1.300

$11,388

Now apply your actual portfolio. Twelve channels in two languages is 24 track-years. On Deepgram streaming that is $97,128 a year. On Azure with language ID it is $273,312. On the cheapest option available it is $31,536, and that number grows every time someone launches a feed.

This is the shape of the problem that sends broadcasters looking at their own hardware. Not because cloud is expensive per unit, but because the unit count only goes up.

Why is raw ASR not the same as captioning?

Before comparing anything, an important correction, because this is where most cost comparisons go wrong.

Every price above buys you a transcript. A stream of words with timestamps. That is an ingredient, not a deliverable. Getting from a transcript to a compliant caption stream requires:

  • Reading rate control. Captions that scroll faster than roughly 180 words per minute are unreadable. Text has to be condensed, not just transcribed.
  • Line breaking and character limits. Typically 32 characters per line, two or three lines, broken at linguistic boundaries rather than mid-phrase.
  • Placement. Captions must not obscure burned-in lower thirds, scores, or tickers. The FCC's caption quality rules at 47 CFR 79.1(j)(2) list placement as one of four explicit standards, alongside accuracy, synchronicity and completeness. Placement is the one automated tools handle worst.
  • Speaker identification. Who is talking, marked so a viewer can follow.
  • Latency inside a regulatory budget. Ofcom's 2024 best practice guidelines ask providers to aim for a mean live subtitle latency no greater than 4.5 seconds across their live programming. Ofcom's earlier review found broadcaster averages of 5.1 to 5.6 seconds. Every processing stage you add spends part of that budget.
  • Packaging. CEA-608/708 for broadcast, IMSC or TTML for streaming, inserted into the right stream with the right timing.

This is why broadcast captioning vendors price differently from ASR APIs. Ai-Media's LEXI Recorded is publicly listed from $0.20 per minute, and the company reports average quality of 98.7% NER for LEXI 3.0, with independent audits finding 35% fewer recognition, formatting and punctuation errors than the previous version. Live LEXI pricing is quoted by sales. So is Enco enCaption, VITAC, Verbit and effectively everyone else in the category. Pricing opacity is the norm in broadcast captioning, which makes honest comparison difficult and is worth naming out loud.

When you do compare, watch the metric. Broadcast vendors quote NER, a weighted model designed for live subtitling that penalises edition, recognition and rendering errors. ASR vendors quote WER on clean benchmark audio. An NER of 98% and a WER of 5% are not two views of the same thing. Ask any vendor for both the metric and the test set, and be suspicious of accuracy claims measured on read speech when your content is a stadium with crowd noise.

While we are correcting things: there is no numeric FCC accuracy threshold. The Commission deliberately declined to set one, judging caption quality against those four standards case by case. The 99% figure everyone quotes is a contract SLA, not a rule.

Human captioning sets the ceiling for context. Professional CART runs from around $70 per hour, and public contract schedules go considerably higher, with one US state listing standard stenographic CART at $108 to $147 per hour and solo long-form assignments at $190 to $240. Broadcast-grade live human captioning reaches $7 per minute. Against those numbers, everything else in this post is inexpensive.

What does self-hosting actually cost?

The models are free. That surprises people.

Model

Avg WER

Throughput

Licence

NVIDIA Parakeet TDT 0.6B v2 (English)

6.05%

RTFx 3,386 at batch 128

CC-BY-4.0

NVIDIA Parakeet TDT 0.6B v3 (25 EU languages)

6.32%

RTFx ~3,333

CC-BY-4.0

Whisper large-v3

7.44%

RTFx 145.5

MIT

Whisper large-v3-turbo

~7.8%

2.7 to 4x large-v3

MIT

NVIDIA Canary-1B-v2 (25 EU languages)

better than large-v3

RTFx 749

CC-BY-4.0

Mistral Voxtral-Mini-3B

7.05%

RTFx 109.9

Apache-2.0

NVIDIA reports Parakeet TDT 0.6B v2 achieving an industry-best 6.05% word error rate on the Hugging Face Open ASR Leaderboard, transcribing 60 minutes of audio in about one second on a single GPU. That is a production-grade open model under a licence permitting commercial use.

Two licensing traps, because they catch people.

CC-BY-4.0 requires attribution. Parakeet, Canary-1B-v2 and Kyutai STT are commercially usable, but if you embed them in a product you owe visible credit. That is a real compliance step, not a formality.

And some models in these same families are non-commercial. The older four-language Canary-1B is CC-BY-NC-4.0. Meta's MMS and SeamlessM4T are non-commercial. Voxtral's TTS release is CC-BY-NC even though its transcription models are Apache-2.0. Read the specific model card, not the family page.

Then there is the managed path. NVIDIA Riva and NIM containers are free for development, but commercial production requires NVIDIA AI Enterprise at $4,500 per GPU per year. For many broadcasters that single line is the reason to serve open models on a plain stack instead.

Hardware, power and space

Street pricing as of August 2026. Note the gap against MSRP, driven by a GDDR and HBM memory shortage that has pushed real prices well above list.

GPU

VRAM

Street price

Power

NVIDIA L4

24 GB

~$2,000 to $2,800

72 W

NVIDIA A10

24 GB

~$3,000 to $3,800

150 W

NVIDIA L40S

48 GB

~$6,300 to $9,000

350 W

RTX 5090

32 GB

~$4,830 against $1,999 MSRP

575 W

RTX PRO 6000 Blackwell

96 GB

~$13,250 against $8,565 MSRP

600 W

Concurrency is what turns hardware into channels. Measured figures using faster-whisper with INT8 quantisation: an L40S sustains 25 or more concurrent Whisper-turbo streams, an RTX 5090 handles about 11 concurrent large-v3 streams, and an H100 exceeds 100 turbo streams. Whisper-turbo at INT8 needs roughly 1.5 GB of VRAM per instance, large-v3 around 2.5 to 2.9 GB.

Treat those as starting points rather than guarantees. Real-time streaming sustains far fewer streams than batch throughput suggests, and your audio is noisier than the benchmark set.

Running costs. US commercial electricity averaged 13.51 cents per kWh in April 2026 per EIA data released 25 June 2026, up 4.8% year over year from 12.89 cents. Germany runs around 27 euro cents for SME bands, the UK 24 to 28 pence, the Nordics closer to 11 euro cents. A GPU node drawing 0.7 kW at the wall consumes 0.7 multiplied by 8,760, which is 6,132 kWh per year. That is about $828 at US commercial rates and roughly $1,656 in Germany.

Rack space is the other line. CBRE's North America Data Center Trends report for H2 2025 put the average asking rate for a 250 to 500 kW wholesale requirement at a record $196.25 per kW per month, up 6.6% year over year, with vacancy at a record low 1.4%. All-in with services runs closer to $250 per kW per month. Power, not floor space, is the binding constraint, and high-density capacity increasingly requires pre-leasing months ahead.

Putting it together for one L40S-class node at roughly $15,000 including the server, depreciated over three years:

  • Hardware: $5,000 per year
  • Power: $828 per year
  • Colocation at 0.7 kW: about $2,100 per year
  • Total: about $7,900 per year

Divided across the channels it serves, that is roughly $790 per channel-year at 10 channels and $395 at 20. Both figures sit below the cheapest streaming API. Both also exclude the thing that actually decides this.

What do the TCO models leave out?

The infrastructure math above is the easy half. The expensive half is people.

Engineering to build it: model serving, stream handling, caption formatting, encoder integration. MLOps and model updates, because new models ship every few months and someone has to evaluate, benchmark and deploy them or you fall behind on accuracy. A 24/7 on-call rota, because a caption dropout on a linear channel is a compliance event rather than a degraded experience. Redundancy, since true N+1 or 2N roughly doubles the infrastructure numbers above. Glossary and custom vocabulary maintenance for player names, place names, sponsor names and show-specific terminology, which is continuous work and is what separates 95% accuracy from 98%. And integration with caption encoders and playout, which is rarely as simple as the datasheet suggests.

For a small broadcaster these can exceed all compute costs combined. Any TCO model that omits them is not a model. It is a sales tool.

Where does the break-even actually sit?

Two thresholds decide this, and neither is the GPU price.

Duty cycle. Self-hosting only beats metered pricing if the capacity stays busy. Against the cheapest batch rates, break-even lands somewhere between 25% and 50% utilisation per track. Below that, metered wins and it is not close. A 24/7 linear channel runs at effectively 100%, which is exactly the workload ownership favours. A department transcribing 100 hours a month is not.

Track count on dedicated hardware. With one appliance serving a handful of channels, fixed cost per channel stays high. Somewhere around eight concurrent tracks the per-channel infrastructure cost drops below the cheapest streaming API, and it keeps falling from there. Below that count an API is genuinely cheaper all-in and you should use one.

General buy-versus-rent guidance from outside our industry agrees. Owning tends to beat cloud GPU rental above roughly 50% to 65% sustained utilisation.

The best-documented example of the pattern is 37signals. CTO David Heinemeier Hansson told The Register in October 2024 that the company cut its cloud bill from $3.2 million a year to $1.3 million on a hardware outlay of around $700,000, without adding staff. That is a stable, predictable, high-volume workload, which is what 24/7 captioning is. The counter-context matters too: global cloud infrastructure spending grew 35% year over year in Q1 2026. Repatriation is a targeted play for predictable workloads, not an industry exodus.

Where does cloud still win?

Be honest about this, because vendors selling on-prem rarely are.

Fewer than about eight concurrent channels, and the fixed costs do not amortise. Bursty or event-driven work, like a tournament three weekends a year, should never be capex. Long-tail languages you caption twice a month do not justify a dedicated model. A one-time archive back-catalogue pass is exactly what Google's Dynamic Batch tier at roughly $0.003 per minute exists for. And if there is nobody to run it, the cheapest infrastructure in the world costs more than an API.

What changes the math beyond cost?

Three factors regularly outweigh the dollars.

Forecastability. Metered pricing means your AI line item moves with programming decisions you do not control. Flat capacity, owned or licensed, gives finance a number that does not move. In a year when cost control is the industry's stated top priority, cited as the leading challenge in Bitmovin's Video Developer Report, with 58.3% of respondents to NewscastStudio's 2026 sentiment survey pointing to both aging infrastructure and pressure to produce more content, predictability has its own value.

Content security and embargo. Sending unaired news, pre-release drama or a pre-feed from a rights-restricted event to a third-party API is a content-security decision, not just a procurement one. For many broadcasters this alone settles the question.

Data residency. Where the audio is processed, and under whose jurisdiction, is now a contractual question in most European deals. Air-gapped and on-prem options exist because of it. Microsoft's disconnected containers start at $74,100 per year for a 120,000-hour block, about $0.62 per hour, which tells you what the market charges for sovereignty.

What should you ask any vendor?

Whether you are evaluating an API, a captioning service or a platform, these four separate marketing from engineering quickly.

  1. Is your pricing metered or flat, and does it change if I self-host? A vendor whose price rises when you move the workload on-prem is charging you for their infrastructure, not their software.
  2. What accuracy metric, on what test set? Insist on NER and WER, measured on audio that resembles yours. Read-speech benchmarks predict nothing about a live sports mix.
  3. What licences cover the models in your stack? If any component is CC-BY-NC you have a problem you may not know about. If it is CC-BY-4.0, attribution is your obligation too.
  4. Where does my audio go, and who can see it? Get the answer in the contract, not the sales deck.

What to do this week

Measure your real duty cycle. Not channel count. Actual hours of audio you need processed per month, per track. This single number decides everything else, and most broadcasters have never calculated it.

Price your own portfolio at three rates. Take $0.15, $0.46 and $1.30 per hour, multiply by your annual track-hours, and put all three in front of finance. The spread is usually the argument.

Benchmark two open models on your own audio. Whisper large-v3-turbo and Parakeet TDT 0.6B v3 both run on a single mid-range GPU and cost nothing to trial. Use your worst audio: the noisy stadium, the accented remote guest, the phone-quality caller. Vendor leaderboards will not tell you what you need to know.

Separate the transcript question from the caption question. Decide what you need delivered, a searchable transcript or a compliant caption stream in a specific packaging format, before comparing any prices. They are different products, and conflating them is the most common budgeting error in this category.

Then look at where the inference needs to sit. If you are already moving toward software-defined production, the same architecture that lets media functions run as containers on shared compute is what makes on-prem inference practical rather than bespoke, and MXL is one of the tap sources an inference function reads frames and audio from. We covered that shift in our guide to MXL and the EBU Dynamic Media Facility.

Sources

What live captioning actually costs in 2026