160–180 wpm
Captions people can actually read: a practical guide for vertical video ads
Set caption timing, line length, placement and contrast for vertical ads using broadcast guidance and accessibility ratios, then check it on a phone.
Your captions have to do more than repeat speech: they have to survive the platform interface, stay up long enough to read, and remain legible against moving footage. This guide gives a vertical-video producer a timing method, a placement check and a contrast test you can apply to an actual ad cut.
Captions carry meaning, not just dialogue#
Some viewers arrive with sound; some cannot hear it; some choose not to use it. A 2016 trade-press story often repeated as “85% watch with sound off” described publisher-reported Facebook News Feed video, not a current cross-platform ad statistic. Meta later said more than 75% of Instagram Reels views were sound on, without publishing the method on that page.[1][2] Neither figure gives you permission to make the story unintelligible in one mode. Make the product and the action visible, and caption the information the audio adds. For the complete beat structure, see the short-form ad anatomy.
WCAG's prerecorded-caption criterion calls for captions for prerecorded audio in synchronized media, subject to its stated text-alternative exception. Its definition includes meaningful sound effects, music, laughter and speaker identity, not just transcribed words.[3] WCAG is a web accessibility standard; this article is not claiming it automatically creates a legal duty for every platform-hosted social ad. It is a useful editorial bar: if a door slam, speaker change or alarm changes the meaning, give the sound a readable text equivalent.
Captions and on-screen copy have different jobs. A caption lets someone follow speech and sound; a headline condenses the proposition; a price and call to action ask for a decision. Do not treat a burned-in slogan as a substitute for captions if the voiceover carries another claim. Nor should you depend on automatic platform captions to rescue text placed underneath the platform's own handle or buttons. Make a reviewable caption track, then inspect the actual viewing mode.
Set the reading pace before animating words#
The BBC's subtitle guidance recommends 160–180 words per minute, roughly 0.33–0.375 seconds per word, and gives about 0.3 seconds per word as a minimum period. Netflix's English SDH guide instead gives an adult ceiling of 20 characters per second; its general guide specifies subtitle events between five-sixths of a second and seven seconds.[4][5][6] These are different production guides, not competing platform mandates. Use them as warning lights, then read the cut at phone size. A slow visual demonstration may need more breathing room than either number suggests.
Take a fictional line, “The filter lifts out without spilling.” That is six words. At the BBC's recommended pace, budget about two seconds or more for the line, then check how long it takes a fresh viewer to locate the relevant object. The rough per-word minimum is not a timer that automatically makes a subtitle readable; punctuation, unfamiliar names, motion and simultaneous labels increase the reading load. If you cannot leave the line up long enough without covering proof, simplify the spoken sentence or split the visual beat, rather than accelerating every letter.
- 1Open: 5 words 0–2s
- 2Demo: 6 words 2–5s
- 3Result: 5 words 5–8s
- 4Ask: 4 words 8–10s
The figure is a fictional timing exercise, not a BBC-approved template. Notice that the demo line has extra room because the viewer must look at the product and read simultaneously. The final ask also has to coexist with the platform's own call-to-action treatment. Keep adjacent captions stable long enough to be read; test transitions by replaying from just before a cut, rather than scrubbing to a still frame. If the speaker races, edit the script or shot length first. Mechanical speed-up cannot make an overloaded frame simple.
For captions tied to a speaker, group words by a complete thought and break at a natural phrase boundary. The BBC favors complete sentences per subtitle and cautions against very short gaps between subtitles.[4] A short ad may require a producer's judgment about where the sentence can be split; preserve its meaning rather than forcing every event to occupy the same duration. Label a sound such as [kettle clicks] only when it communicates something a viewer otherwise misses. Do not caption decorative music repeatedly if it adds no new meaning; do identify a relevant change in music or the arrival of another speaker.[3]
Design lines for a narrow frame#
Landscape subtitle conventions do not fit a phone held upright. The BBC's guidance equates its vertical online width with about 25 characters per line inside 90% of a 9:16 frame, and recommends at most three lines in vertical video. Netflix's English guide allows 42 characters per line and a maximum of two lines, designed for a different delivery context.[4][5] Neither character count is a universal law for a particular font. A wide capital or non-Latin script changes what fits; measure the rendered line, not only its length in a text editor.
| Production reference | Width and lines | How to use it on a vertical ad |
|---|---|---|
| BBC online vertical | About 25 characters per line in a 90%-width region; up to three lines[4] | A useful starting envelope before checking platform overlays |
| BBC broadcast landscape | 37 characters per line, typically two lines[4] | Do not paste the line breaks unchanged into portrait |
| Netflix English timed text | 42 characters per line, at most two lines[5] | A delivery-specific specification, not a portrait safe-area template |
There is a small inconsistency inside the BBC document: one landscape-width table entry says 68%, while its explanatory equivalence uses 75%. The vertical guidance—90% width and about 25 characters—is consistent in both places.[4] Keep the vertical number as a starting point and test the actual font. If a product name pushes a line across the side of the frame, shorten the adjacent words. If a line wraps between two visual ideas, rewrite it so the caption follows the cut rather than fighting it.
A third line is an allowance, not a target. When three lines cover the product, use two or shorten the narration. Give the captions their own hierarchy: stable body weight, restrained emphasis and line breaks that help comprehension. All-caps, one-word-at-a-time karaoke and simultaneous animated headlines can each compete with the demonstration. The reader should know where to look without learning a new typographic system every shot. Watch one complete pass without pausing and write down any word you missed; those are edit points.
Put the words where the interface cannot eat them#
At export, the video is a rectangle. In the feed, it becomes a rectangle behind controls. Meta recommends leaving at least 14% at the top, 35% at the bottom and 6% on either side of a Reels ad free of key text and logos. That is approximately 269 top, 672 bottom and 65 side pixels on a 1080 × 1920 working frame—our arithmetic, not a Meta pixel specification.[7] The figure shades those margins. It does not claim that every viewer's controls have identical dimensions.
The shaded central region is a placement boundary, not a command to center every subtitle vertically. Move the line high enough to avoid the lower caption and CTA controls but low enough not to cover the speaker's eyes or the action. If the background changes rapidly, use a solid caption box or a tracked patch of calm footage. Check the ad in the platform's preview and, where possible, on a real phone. Meta says a disclaimer-bearing Reels ad needs a larger 40% bottom clearance, and it may zoom or letterbox some taller screens.[8] The same export can therefore have more than one credible crop.
TikTok says its safe area depends on caption length and additions to the ad; YouTube supplies a separate vertical overlay. Their numbers and the shared box are worked through in the TikTok, Reels and Shorts safe-zone guide.[9][10] Do not assume that the caption position that passed one placement will survive another. When the ad carries a material-connection disclosure, keep that text legible as well: the FTC looks at contrast, position and screen time, and its examples call out inadequate disclosures on phones.[11]
Make the actual text-background pair readable#
WCAG 2.2 specifies 4.5:1 contrast for ordinary text and 3:1 for large text under its Level AA minimum-contrast criterion. Its formula compares the relative luminance of the lighter and darker colors.[3] Burned-in text within a larger image has an interpretive exception in the standard; as a production choice, treat captions as readable text rather than using that exception as an excuse for low contrast. The background in the calculation is the actual pixel behind the letters, or the solid caption box when the box is opaque.
The figure computes the ratios from the color values using WCAG's formula; the color choices are illustrative, not reported platform data. A 60% opaque black overlay on pure white footage becomes #666666: our derived color calculation, because 40% of each white channel remains visible. White text on that resulting gray has roughly 5.74:1 contrast, compared with 21:1 on opaque black.[3] Over darker footage, that particular black overlay makes a darker box; over colored footage, sample the actual composited pixels anyway. If a caption box has another opacity, color, gradient or blend mode, the white-footage calculation no longer describes it. An outline alone can fail where the letter interiors touch a busy light patch. Verify the finished export, because compression and resizing can change thin strokes.
Do a worst-frame pass rather than taking one attractive screenshot. Pause on the lightest point of the shot, the transition into the next shot, and any frame where a moving highlight crosses a word. If the ratio fails there, increasing font weight alone may not solve it: first give the letters a reliable background, then recheck the crop. A price or legal label deserves the same attention as a caption; a line that viewers cannot read does not become useful because the audio says something similar. Record the tested foreground and composited background colors with the export so another editor can reproduce the decision.[3]
Avoid using hue alone to identify the speaker or distinguish an offer from a disclaimer. Add a speaker name, position or wording cue where needed. Large type can qualify for the lower contrast threshold under WCAG's definition, but a style that looks large on an editor monitor may be small on a phone.[3] Judge the actual display size rather than choosing the easier ratio because the design file calls a style “large.”
Motion, flashes and automatic captions#
Word animation is a pacing decision. A flash or sudden full-frame inversion can make an ad uncomfortable or inaccessible. WCAG's Three Flashes or Below Threshold criterion limits flashing to no more than three times in any one-second period unless the general and red-flash thresholds are met; the criterion applies to web pages containing video too.[3] It is safer to avoid flashing as a hook than to try to prove a complicated threshold after export. Use a cut, movement of the product, or a stable typographic change instead.
Automatic captions are useful as a draft. YouTube says automatic captions can be published on Shorts when available and explicitly asks creators to review and correct them. TikTok says uploads can get generated captions and lets creators edit or remove them.[12][13] Check names, product terms, numbers, negations, punctuation and speaker changes against the audio. “Does not contain” becoming “does contain” changes a claim, not just a typo. If you burn captions into the video and the platform also displays its own, preview for duplicate overlapping lines and decide which text track is appropriate for the upload.
The speech transcript is not a complete accessibility pass. A viewer who cannot hear the click that proves a latch closed needs that information in the caption or in a visible close-up. WCAG's audio-description criterion also addresses important visual information missing from the soundtrack in prerecorded synchronized media; W3C notes that separate description is unnecessary when the existing audio already explains all important visuals.[14] A concise voiceover describing the action can therefore serve people listening without watching, while the captions serve people watching without sound. Check the cut both ways.
Worked example: caption a fictional product demonstration#
Here is a fictional twelve-second ad for a reusable filter. The timings and script are editorial examples, not platform metrics. The product is shown during the claim about removing it; the narrator does not make an unsupported performance promise. Each line fits the narrow frame when rendered in the chosen font; the durations leave a viewer time to inspect the action.
| Time in fictional cut | Picture and sound | Caption event | Editorial check |
|---|---|---|---|
| 0–2.5 s | Hand lifts filter from a cup; narrator begins | “This filter lifts out cleanly.” | Five words; keep the hand and line apart |
| 2.5–5.5 s | Close-up of the rim; audible click | “The rim clicks into place.” / “[click]” if not visually clear | Do not crowd the useful close-up with extra slogans |
| 5.5–8.5 s | Water is poured through the real product | “Rinse it, then use it again.” | Shows the demonstrated step; no invented savings claim |
| 8.5–12 s | Pack and clear next action | “See how it fits your cup.” | Leave the ask readable beside the destination's own controls |
Read it once with the screen covered. If the narration fails to tell you what happened visually, revise it; if the spoken version makes an unverified claim, revise the claim. Then watch muted and see whether the click and purpose are still clear. Finally check the line in the exported video at phone size over the brightest frame and in the destination preview. A caption that looks excellent on a paused editor canvas can disappear behind a live CTA.
Common mistakes#
Equating captions with a headline. A text hook does not communicate the rest of a voiced demonstration or a meaningful sound. Keep a caption track that follows the audio, and separately decide which on-screen words carry the offer.[3]
Applying a landscape line limit as a vertical target. Netflix's English line limit and the BBC's portrait guidance describe different frames and workflows. Reflow and inspect your actual typeface rather than citing one number as universal.[4][5]
Timing to the editor's reading speed. You have heard the line dozens of times. Show the cut cold to somebody unfamiliar with the product; shorten the script where their eyes must move between caption and proof. The BBC speed is a starting measure, not a substitute for this pass.[4]
Testing contrast on an empty background. Moving footage changes behind translucent text. Calculate at the weakest composited frame and check phone-size strokes.[3]
Assuming the auto transcript is final. Automatic captions can be wrong about a brand name, number or negation. Review the published caption state, not just your local export.[12][13]
Checklist#
- Make the ad understandable with sound on and off; caption speech, important sound and speaker changes.[3]
- Split lines at natural phrases, set readable durations, and check vertical width in the rendered font.[4]
- Place every essential word inside the destination's current safe region and inspect the live controls.[7][9]
- Measure contrast against the brightest composited background, or use a sufficiently opaque box.[3]
- Remove unnecessary flashes and verify the whole exported cut rather than a paused still.[3]
- Correct generated captions, check duplicate overlays, and watch the upload at phone size.[12][13]
Keep reading
Safe zones for TikTok, Reels and Shorts: where your text actually survives
Compare official TikTok, Reels and Shorts safe-area templates in pixels, find the shared box for text, and understand why published measurements differ.
AI avatars, UGC-style ads and the rules: what to make and disclose
Follow the disclosure decisions for actors, AI avatars, cloned voices and customer claims across FTC guidance and platform labels, with scenarios.