Meet Amira Labs at IBC 2026. Hall 1, booth 1.C37k. RAI Amsterdam, 11–14 September.

Live caption QC: why a validator pass is not FCC compliance

Live caption QC: why a validator pass is not FCC compliance

Kyle Suess

The caption stream is present. The file validates. The transport monitor is green. On the viewer's screen, the captions cover the score and finish the commentator's sentence after the next play has started.

That is the gap live caption QC has to close. Technical validation can establish that caption data meets particular format and delivery checks. It cannot, by itself, establish that a deaf or hard-of-hearing viewer received an understandable version of the programme.

For an engineering team, the practical consequence is straightforward: test the delivered picture, audio and captions together. For the person buying a monitoring system, ask exactly which of those components its pass indicator covers.

![A fictional broadcast operator reviewing caption bands against weather, sports and news pictures](https://media.amiralabs.com/blog/live-caption-qc-fcc-quality-latency-placement-2026/4a80402d5372-live-caption-qc-hero-1600x838.jpg)

AI-generated editorial illustration. The screens and facility are fictional, not an Amira product interface or a real installation.

What does a caption validator actually prove?

A validator proves that the checks it implements passed for the input it inspected. The useful question is what remains outside that inspection.

Separate the evidence into layers. A presence check establishes whether expected caption data is arriving. A format check examines the applicable syntax and profile. A rendering check examines the displayed result. An editorial review compares that result with the programme's meaning.

These layers need different inputs. A file-only check cannot inspect a score graphic added downstream. A successful test before packaging cannot demonstrate what every receiving application displayed. A transcript comparison without the picture cannot decide whether captions obscured a speaker's face.

There is a current standards reason to be precise about the format layer. W3C published IMSC Text Profile 1.3 as a Recommendation on May 21, 2026. It updates the earlier text profile, including support for superscript and subscript, and guidance for Japanese authoring. Its publication does not demonstrate support in an installed receiver estate. W3C's release announcement.

For delivery chains using IMSC, specify the required profile and supported receiver behaviour in the test plan. The specification distinguishes document and processor conformance; its referenced Hypothetical Render Model addresses document complexity and rendering at authored display times. Those are valuable technical checks. They do not establish whether an authored sentence accurately represents the speaker. IMSC Text Profile 1.3, sections 5 and 8.2.

Which FCC caption-quality standards should the test plan cover?

The four categories are accuracy, synchronicity, completeness and placement. The FCC's quality framework concerns covered US television programming; do not assume that every internet video has identical obligations. Establish the service's applicable rules with the compliance team before turning requirements into acceptance tests. FCC television-captioning compliance guide.

The table below translates those categories into suggested engineering checks. It is a test-design aid, not an FCC certification checklist.

Quality category

What the category addresses

Suggested automated signal

Suggested human review

Accuracy

Speech and relevant non-speech information are represented correctly.

Flag differences against a reviewed reference; locate unexpected names and numbers.

Check changed meaning, missing speaker context and significant sounds.

Synchronicity

Captions follow the corresponding audio and remain readable.

Measure matched speech-to-caption lag and display duration.

Watch rapid exchanges and changes of speaker.

Completeness

Caption coverage extends through the programme.

Flag unexplained gaps and suspicious start or end behaviour.

Check the opening words and the final sentence before a break.

Placement

Captions remain visible without obstructing essential picture information.

Flag overlaps with defined graphics regions or clipped text.

Inspect faces, scores, maps and changing graphics in context.

The underlying categories are set out in the FCC guide's quality-standards section. The proposed signals are operational recommendations; no individual signal establishes compliance.

Live and near-live programmes receive a circumstances-based assessment. The FCC considers their production difficulties, overall understandability, technically feasible efforts to reduce lag, transition cut-offs and susceptibility to unintended obstruction. It also recognises de minimis errors. A single typo is therefore not automatically a violation; live production is not a blanket exemption either. FCC guide, application of standards.

The rule text and guidance discussed here were checked on August 30, 2026. The established FCC quality framework predates 2026; the IMSC revision is a separate technical development. For the broader jurisdictional picture, see our broadcast caption-compliance guide.

Does the FCC require 99% caption accuracy?

The FCC television-caption quality rule does not establish 99% as a universal minimum accuracy threshold. Be careful, however, with the opposite claim that the number never appears in the rules.

Section 79.1(k)(2)(iv) uses 99% in a worked example of an accuracy calculation: 7,000 programme words, less 70 errors, produces 6,930 correct words; dividing by 7,000 produces 99%. The vendor best practices call for metrics, minimum acceptable standards and regular evaluations. The worked example does not set a universal passing score. The FCC's own guide reproduces that distinction. FCC guide, real-time captioning vendor best practices.

A contract can set a numerical target. Its usefulness depends on the scoring method, reference transcript, sampling and treatment of errors. Ask the supplier to explain each before comparing percentages.

Consider an invented example: a caption drops “not” from “the bridge is not open.” One missing word reverses the message. A programme-wide percentage can make that event look small even when a viewer would consider it consequential.

Keep the aggregate score, but retain the examples behind it. Review names, numbers, negation and attribution as distinct error categories. And keep accuracy separate from placement: a perfectly transcribed sentence can still cover the information a viewer needs.

This also changes model evaluation. The broadcast speech-recognition model-selection guide is an upstream decision. Caption QC must evaluate what survives the entire delivery path.

How should a broadcast team measure live caption latency?

Measure the same spoken event and its matching caption at explicitly identified observation points. A latency number without those endpoints is difficult to interpret or reproduce.

Keep two measurements separate. The FCC's live-caption vendor best practices describe lag from speech supplied at the programme origination point to captions received at that same point. Receiver-side measurement answers an additional question: how far behind the audio does the viewer see the caption? 47 CFR §79.1(k)(2)(vi).

For an illustrative receiver test, use one common source event and a shared time reference:

  • The matching audio is heard at the test receiver 2.0 seconds after the source event.
  • Its caption becomes visible there 6.2 seconds after the source event.
  • The receiver-side caption lag is 6.2 − 2.0 = 4.2 seconds.

These are chosen example values, not a measured service result or an acceptance threshold. Measuring only the caption generator would miss the difference between the audio and caption delivery paths.

![Illustrative timing diagram: audio reaches a test receiver at 2.0 seconds and the matching caption at 6.2 seconds, yielding 4.2 seconds of viewer-side lag](https://media.amiralabs.com/blog/live-caption-qc-fcc-quality-latency-placement-2026/6bdf901f5cf2-caption-lag-measurement-1600x1000.png)

Original Amira Labs diagram using hypothetical values. Both observations concern the same event at the same receiver. This viewer-side test is additional to the FCC best-practice origination-point measurement.

For international distribution, label the jurisdiction beside the metric. Ofcom's access-services guidelines, paragraph 3.9, say providers should aim for a mean latency of no more than 4.5 seconds across live programming taken together. That is UK guidance, not an FCC limit, and not a maximum for every individual caption. Ofcom access-services guidelines.

A single 4.2-second observation cannot establish whether that mean target is met. Define the sample and measurement method. For engineering diagnosis, also track the median, slow-tail behaviour and maximum observed lag, with sample counts and programme context. Those additional statistics are suggested diagnostics, not newly invented regulatory thresholds.

Also record whether the measured display contains provisional text or a stable caption. A fast first word and a readable, completed caption answer different questions. Any decision to delay programme audio and video to improve alignment needs an approved production delay budget and an end-to-end test.

Why does caption placement require the actual programme picture?

Placement quality depends on what occupies the screen at that moment. A valid coordinate inside the image can still be an unsuitable location for a caption.

The FCC placement standard addresses interference with essential visuals, including faces and featured graphics, as well as overlapping, off-screen or illegible captions. 47 CFR §79.1(j)(2)(iv).

For a practical test, use representative output frames with the graphics package enabled. Include a scoreboard, a map, a two-person interview and a busy lower third. Repeat the review on supported receiver types and relevant caption-size settings. Record the tested device and configuration so the result can be reproduced.

![Paired fictional sports frames show a caption obstructing the lower score graphic, then moved upward to leave the score visible for this shot](https://media.amiralabs.com/blog/live-caption-qc-fcc-quality-latency-placement-2026/0885fc176d4b-caption-placement-review-1600x1000.png)

Original illustrative layout. Moving a caption upward works for this particular frame; it is not a universal safe-position rule.

An overlap detector can identify a collision with a known graphics region. Treat that as a review prompt. The significance of the collision still depends on the content: covering empty grass is different from covering the score, even when the affected areas are equal.

Avoid fixing every problem with a permanent top-of-screen position. The next shot may put a face or essential graphic there. Test the placement policy across a sequence, including the transition between layouts.

How can caption-presence checks miss incomplete coverage?

A presence indicator tells you that data exists at the observed point. Completeness testing asks whether the relevant programme content was represented throughout the interval being reviewed.

Build a test sequence with known opening words, a rapid exchange, a significant off-screen sound and a sentence that finishes immediately before a programme break. Review the decoded output against that sequence. Check what happens after a caption-source switch and when the programme resumes.

These tests can expose several different faults: old text remaining on screen, delayed captions disappearing at a transition, or a language service carrying the wrong material. Give each fault its own label so the operator can route it to the right owner.

Do not turn every speech-detector gap into a confirmed caption failure. A detector may be wrong; the source may contain silence, unintelligible speech or music. Conversely, speech detection alone will not establish whether a meaningful non-speech event needs description. Review suspicious intervals with the source audio available.

The useful output is a timestamped discrepancy with context, rather than an unexplained red badge. A short excerpt around the event, where recording is permitted, is much easier for the caption provider or distribution team to investigate.

What should master control receive when QC finds a problem?

Master control should receive an actionable incident with a named owner, the affected output and a clear recovery path. Another undifferentiated alarm adds work without resolving the caption problem.

For each event, capture the programme, language service, observation point, time reference and detected symptom. Distinguish loss of data from late data, and both from a suspected editorial error. Include the relevant software or receiver configuration when it could affect reproduction.

Agree on escalation before the shift starts. Engineering can investigate delivery and rendering faults. The caption provider can investigate source-text or captioner problems. Production can review graphics conflicts. The compliance team can assess applicability and recurring issues. A routing plan should identify actual contacts rather than assume all four functions are staffed continuously.

Keep a tested fallback in the runbook. Exercise it on a controlled test path or in an approved maintenance window. Confirm that switching sources restores useful captions and does not merely restore the presence indicator.

There is a specific recordkeeping distinction worth preserving. Section 79.1(c)(3) requires distributors to retain monitoring and maintenance records, including technical equipment checks, for at least two years. That provision does not itself require keeping two years of every programme recording. 47 CFR §79.1(c).

Set retention for review clips and transcripts separately with the appropriate rights, privacy and compliance owners. Keep enough linked evidence to explain the incident and corrective action without treating unlimited recording as the default.

Where does automated caption QC stop being reliable?

Automated QC becomes weak evidence when its reference is unverified or its inspection point excludes the failure being assessed.

A second speech recogniser can help locate possible transcription discrepancies. It remains another fallible transcription system. Agreement between two systems is not proof that either captured a proper name, number or ambiguous phrase correctly. Use reviewed references for acceptance testing and preserve uncertain cases for human assessment.

Picture analysis has similar boundaries. Detecting a face or a block of text does not by itself establish what information matters to the viewer. A monitoring tool should disclose its tested languages, layouts, inputs and failure cases, and provide a way to inspect the evidence behind a flag.

Include reviewers who rely on captions when designing representative tests. Ask whether they can follow speaker changes, simultaneous discussion and important sound cues. A hearing reviewer who already knows the dialogue can unconsciously fill in omissions.

This is where procurement questions become concrete. Ask a supplier to demonstrate a technically valid but mistimed caption, a graphics collision, a missing final sentence and a meaning-changing transcription error. For each, ask what it detects, what it misses and what reaches the operator. Treat untested cases as unknown.

What can your team improve this week?

Start with one programme and one delivered output, then make the evidence repeatable.

  1. Map the observation points. Identify where you check presence, format, origination-point lag and receiver-side presentation. Name the gaps.
  2. Build a small challenge sequence. Include a graphics conflict, a critical word, a programme transition and a caption-source recovery. Use material you are authorised to test.
  3. Write acceptance criteria and ownership together. State the metric, method, programme type and escalation contact. Label contractual targets and internal alert settings separately from legal requirements.
  4. Review the result with a caption user and an operator. Keep the failures that the automated report missed, and add them to the next test.

The next green QC report should tell the shift what was checked, where it was checked and which parts still need review. That is an evidence trail an engineering team can use and a buyer can evaluate.

This article provides engineering guidance, not a legal determination of compliance. Confirm the rules applicable to each service and jurisdiction with your compliance advisers.

Sources