Live vertical sports video: keep the action, score and sponsor in frame

Live vertical sports video: keep the action, score and sponsor in frame

Kyle Suess

A fixed centre crop from a 1920 × 1080 sports feed to 9:16 keeps a strip only about 608 pixels wide. That is 31.6% of the original picture width. The other 68.4% disappears.

That geometry is why live vertical sports video cannot be accepted by watching whether an athlete remains inside a tracking box. The crop may contain the ball carrier and still lose the defender, the finish line, the official, the scorebug, the caption, or the sponsor element that paid to be visible. A technically valid 9:16 stream can tell the wrong story.

The category is moving quickly before IBC 2026. FOR-A announced Go Vertical! AiDi as a dedicated product on September 1, AWS made Elemental Inference generally available in February, and YouTube now supports simultaneous horizontal and vertical live streams. Buyers have more ways to create vertical output. They now need a disciplined way to approve it.

![Conceptual 16:9 sports feed passing through a reframing stage into a verified 9:16 mobile output](https://media.amiralabs.com/blog/live-vertical-sports-video-quality-control/362f23d42aaa-live-vertical-sports-video-quality-control-hero-1600x838.jpg)

Why is vertical sports video now a live-output problem?

Vertical sports video is becoming a separate live service, not merely a social edit made after the event.

YouTube's current live-streaming guidance lets a channel send 16:9 and 9:16 versions at the same time. Viewers in the Shorts feed see the vertical version while other viewers see the horizontal stream. YouTube also offers centre crop, fit-to-phone-width and stacked layouts in its first-party workflow. That creates a real operational fork: two audience experiences can leave the same production at once.

The vendor activity points in the same direction. FOR-A's September 1 announcement says Go Vertical! AiDi creates real-time 9:16 output from existing 16:9 programming and supports automated tracking modes for basketball, soccer and baseball. Nippon TV says the system will be demonstrated at IBC, September 11–14, at FOR-A's stand 2.B53. AWS says Elemental Inference can create vertical versions and highlight clips while live video is being encoded. These are vendor descriptions, not independent acceptance results.

The audience proposition is already visible in production. Sports Video Group reported in April that Peacock planned an AI-driven 9:16 feed for an NBA game inside its mobile Courtside Live experience. AWS described a Fox Sports workflow that detects moments, produces vertical crops and sends clips to a review portal within seconds. Those deployments show that the question has advanced beyond whether automated reframing is possible.

The harder question is whether each delivered version remains editorially and commercially correct.

What must a vertical sports feed preserve?

A vertical feed must preserve the event state a viewer needs to understand the moment, plus every required element that travels with that version.

Write that requirement as an output contract before testing any product. A useful contract names the destination, sport, programme type, required graphics, permitted delay and fallback. It should also say who can approve exceptions. Without that document, reviewers tend to score framing by feel.

Output element

What “present” should mean

Failure example

Evidence to retain

Primary action

The decisive subject and object remain visible through the play

Ball leaves frame while the tracked player remains centred

Timecoded 16:9 and 9:16 comparison

Game context

Enough surrounding action remains to understand cause and consequence

Receiver is visible but defender and boundary are lost

Before, during and after frames

Score and clock

Correct game state is readable when the format requires it

Scorebug is cropped, scaled too small or covered by app controls

Device capture plus reference feed

Captions

Text is complete, synchronized and legible in the destination player

Burned text is clipped or overlaps a platform overlay

Caption sample and device capture

Sponsor element

Contract-required placement survives composition and display

virtual board or sponsored graphic falls outside the crop

Approved placement reference and observed output

Editorial transition

Cuts, replays and interviews use an appropriate composition

Crop chases the outgoing shot after the director cuts

Timecoded transition record

Recovery

The service returns to an approved layout when confidence or input fails

Frozen crop persists after tracking loss

Alarm, action and recovery timestamps

“Keep the athlete centred” is too narrow. In basketball, a crop may need the ball handler, defender and rim. In tennis, it may need both players and enough court to read the rally. In motorsport, a close crop of one car can remove the overtake. Each sport has its own visual grammar, and each programme can contain several shot types.

Why does good object tracking still produce a bad crop?

Object tracking answers where a detected subject moved. A production crop must also decide which subject matters, how much context belongs around it and when the shot's meaning changed.

Consider a football pass. The quarterback begins as the obvious subject. The ball then becomes the link between two areas of the wide frame. A crop that waits for the receiver to become dominant can arrive late. A crop that pans aggressively may be hard to watch. A crop that zooms out far enough to protect the play may leave important objects too small on a phone.

Cuts make the problem sharper. A wide game camera, tight player shot, replay, crowd reaction and interview need different rules. Motion across a cut can resemble object movement even though the director has started a new visual sentence. Occlusion adds another failure mode: a player passes behind an official, leaves the frame, or merges into a group. FOR-A says its product can resume tracking when a registered subject temporarily disappears. That is a useful capability to test, but the buyer still has to judge the composition before, during and after recovery.

Sport-specific tuning matters. Fox Sports told Streaming Media that an early version of its AWS-assisted output was “a little choppy, a little rough” and improved during subsequent work. AWS says Fox and AWS refined the approach over 18 months. That is vendor and participant reporting, yet it gives buyers the right expectation: commissioning requires representative material, production judgment and iteration.

The acceptance target is not a perfect tracking score. It is an output a producer would allow an audience to see.

How should you test automated sports reframing?

Test automated sports reframing on a shadow output with a planned sequence of difficult moments, then score the delivered 9:16 version against the output contract.

Start with recordings you are permitted to use. Include full programmes, not a highlight reel of clean examples. The difficult material often sits between headline moments: substitutions, stoppages, graphics changes, crowd shots, ceremonies, weather delays and returns from replay.

![Live vertical sports video acceptance matrix covering content, graphics, delivery and recovery](https://media.amiralabs.com/blog/live-vertical-sports-video-quality-control/25b586012a9c-vertical-sports-output-qc-matrix-1600x1050.png)

Build the test set across four layers:

  1. Content. Fast direction changes, multiple plausible subjects, small or hidden balls, player clusters, officials crossing the action, celebrations and off-ball incidents.
  2. Production. Hard cuts, dissolves, replay wipes, split screens, picture-in-picture, studio inserts, interviews and a sudden return to live play.
  3. Graphics and audio-linked text. Scorebug changes, lower thirds, lineups, statistics, captions, translated text and sponsored overlays.
  4. Delivery. Each target app, representative phones, player controls, live chat or commerce overlays, orientation changes, bitrate shifts and an unavailable reframing worker.

For each event, retain the reference input, vertical output, crop coordinates if available, destination capture and operator action. Mark the first bad frame and the first good frame after recovery. Count failures per event type rather than collapsing everything into one accuracy percentage. A system can score highly on quiet wide shots and fail every replay transition.

Use a producer and an operator in the review. The producer judges story and composition. The operator judges alarms, evidence and recovery. Add commercial or accessibility review where the output contract requires it.

How do scorebugs, captions and sponsors fail on a phone?

Scorebugs, captions and sponsor elements fail because the video crop and the destination interface reserve different parts of the same small screen.

There is no permanent universal safe rectangle for every vertical platform. TikTok's June 2026 in-feed specification says its safe zone changes with the aspect ratio, caption length and additional formats, and warns that previews can differ slightly from the live version because the preview is not device-specific. Google's YouTube ad guidance says overlays, calls to action and buttons can appear in different positions by format, campaign and screen. It also says the vertical player can compress toward 1:1 during some interactions.

Treat platform templates as inputs to testing, not as a timeless master. Record the platform and app version used for each device capture, and repeat the check when the interface changes.

Graphics need version-aware composition. A scorebug burned into the far corner of the 16:9 master may disappear in a centre crop. Shrinking the entire 16:9 picture inside 9:16 protects the graphic but makes the game image smaller. Rebuilding the scorebug for the vertical version gives better control, although it creates another graphics output that must receive the correct data and survive failure.

Captions need the same distinction. Platform-rendered captions can move within the player and may respond to user settings. Burned-in captions cannot. If burned text is required, validate line length, placement and contrast on the actual destination. If a platform renders the captions, verify that the right timed-text track reaches that version and remains synchronized.

Sponsor treatment is contractual, not inferential. A tracking system should not decide which placement is required. Give the test team an approved list for that output and confirm visibility across live action, replay and interface overlays. Proof that a logo appeared in the encoded video is still different from accredited impression, attention or sponsorship measurement.

How much delay can vertical processing add?

The permitted delay depends on how the vertical version will be used, so there is no honest universal target.

A full-game stream used alongside a television feed may need close enough alignment to avoid spoiling the play. A mobile-only stream can tolerate a different delay if viewers do not compare it with another source. A social highlight can arrive later and still meet its purpose. Write the use case before choosing the threshold.

Measure delay from a shared reference. Put a visible and audible marker into the controlled test feed, capture the 16:9 and 9:16 outputs, and compare the same event. Separate ingest-to-output processing time from player buffering and network delivery. Record median and worst observed results over a full event, including scene changes and recovery.

Vendor figures illustrate why the measurement boundary matters. AWS currently advertises 6–10 seconds for Elemental Inference. FOR-A describes its on-device approach as near real time and says it avoids the five-to-ten-second delay associated with some competing methods. These statements use different products and may use different test boundaries. Neither should become your SLA without a test on your signal path.

Also measure crop motion. A low-latency output that snaps, oscillates or arrives at the subject after the decisive moment is not successful. Timing and composition belong in the same acceptance report.

When is automated reframing the wrong choice?

Automated reframing is the wrong default when the programme's meaning depends on wide spatial context, the destination volume does not justify commissioning, or a missed element carries high editorial or contractual risk.

Some sports are more forgiving than others. A close contest around one basket offers clearer focal points than a formation spread across an entire pitch. A podium ceremony, injury, protest or disputed officiating moment may require deliberate human framing. A premium sponsor activation may need a producer-approved shot rather than a crop inferred from player motion.

Low volume changes the economics. If a team produces a handful of high-value vertical events, a dedicated operator or vertical camera may be simpler and more controllable than building an automated workflow. Automation becomes more compelling when many simultaneous events or clips make manual versioning impossible, but scale does not remove the need for acceptance criteria.

Keep a fallback. That may be a wider crop, a fit-to-width layout, the original 16:9 feed within a vertical canvas, a static safe composition, or withdrawal of the vertical output. The right option depends on the destination. The operator should know which fallback is active and why.

This is where content-aware monitoring matters. Amira Labs approaches live-media intelligence as evidence for an operator decision. In a vertical workflow, that means observing the version that leaves for the audience, not assuming the reframing stage produced what its control panel intended. The acceptance principles are the same ones in our AI broadcast monitoring buyer's checklist: inject known faults, retain timecoded evidence and test failure recovery.

What should you ask a vertical-video vendor at IBC?

Ask for a difficult, repeatable test rather than another highlight montage.

  • Which sports, shot types and programme elements have separate logic, and which fall back to generic tracking?
  • Can we compare the 16:9 input, crop decision and delivered 9:16 output on the same clock?
  • What happens across cuts, replay transitions, occlusion and loss of the tracked subject?
  • How are scorebugs, captions and required sponsor elements protected or recomposed?
  • Can an operator override the crop, set a safe layout and see when automation confidence falls?
  • What is the measured delay boundary, and does the number include encoding, packaging and player delivery?
  • What happens when the model, network or graphics data is unavailable?
  • Can we export timecoded evidence from a full-event acceptance run?

If a system can change a live output automatically, define its permission and rollback boundaries as carefully as its visual tests. Our guide to human approval for agentic broadcast AI provides a practical starting point.

What can your team do this week?

First, calculate the crop. Put a 608-pixel-wide guide over a 1920 × 1080 recording and watch ten minutes of real programme transitions. The failures become obvious quickly.

Second, write one output contract for one destination. Name the required action, graphics, captions, sponsor treatment, delay and fallback.

Third, assemble a twenty-event challenge reel from cleared material, then run it through the candidate workflow and capture the actual destination on at least two representative phones.

Fourth, bring the worst five cases to IBC. Ask each vendor to show the input, decision, result and recovery. A good demonstration should make failure visible, because that is how you learn whether the service is safe to operate.

Sources

Live vertical sports video: keep the action, score and sponsor in frame