Captions on short-form video are not an accessibility addition placed over a finished frame. They occupy a large part of the visible area and are read first, which makes them a structural element of the shot.

Most viewing starts without sound

Feeds commonly begin playback muted or at low volume, and viewers frequently watch in places where audio is impractical.

Under those conditions the text is the only content carrying the words, and a video without it delivers nothing until the viewer chooses to enable sound.

Since that decision is made within the first seconds, text is what determines whether audio is ever turned on at all.

Text is processed before imagery

Readers encountering words in a frame read them involuntarily, and that attention is taken from whatever else is in the shot at the same moment.

Placing text over the subject therefore competes with the subject rather than supporting it, and the viewer resolves the conflict by reading.

Composition has to allow for this by leaving the text a defined region and keeping the important part of the image outside it.

Interface elements claim large parts of the frame

Vertical feeds overlay account names, captions, buttons and progress indicators, and those regions differ between services and change with updates.

Text placed in the lower portion of the frame is routinely obscured, and text at the very top can be cut by device hardware.

Working within a safe central band is therefore not conservatism but a requirement, since anything outside it may not be visible to a substantial share of viewers.

Pacing of text is a timing decision

Words appearing in step with speech hold attention, while text that lags or leads creates a mismatch the viewer notices without identifying.

Blocks that change too rapidly cannot be read, and blocks that persist too long invite the viewer to look away, so the rhythm is being edited alongside the picture.

Automatically generated captions rarely get this right, since they follow speech timing without regard to reading speed or to where the emphasis falls.

Style choices carry more weight than they appear to

Typeface, weight and outline determine legibility against moving footage, and text that works over a static background disappears over a busy one.

Heavy weights with clear separation from the image remain readable on small screens in poor conditions, which is where most of this material is actually watched.

Because the text is on screen for most of the video, these choices define the visual character of the piece as much as any decision about the footage does.