top of page

How to Make Videos Accessible: A Practical Guide

You've just received an accessibility audit notice. The results are uncomfortable: a large lecture library has been uploaded to the LMS, but many recordings still lack properly checked captions, transcripts, or descriptions of essential visual information. The teaching team now has to support current learners while repairing years of content, often without taking every video offline.


That situation is common because video accessibility isn't a single export setting. It's a production and publishing workflow involving recording quality, caption review, transcripts, player controls, LMS embeds, screen readers, and ongoing maintenance. The practical question isn't only how to make videos accessible once. It's how to make accessibility repeatable across hundreds of recordings.


Why Video Accessibility Matters More Than You Think


Accessibility failures affect learning before anyone files a complaint. A student who can't hear a lecture without captions, follow a diagram without spoken explanation, or search a recording without a transcript has to spend effort decoding the format instead of learning the subject.


UK public-sector requirements became clearer when the Public Sector Bodies (Websites and Mobile Applications) Accessibility Regulations applied to time-based media published on or after 23 September 2020. Government guidance expects appropriate alternatives, including captions and transcripts, and says captions should be synchronised with the audio and checked for accuracy rather than published unchecked. See the UK Government guidance on accessible audio and video for the production implications.


The audience is broad. The government's Video-on-Demand Accessibility Impact Assessment estimates that about 5.8 million people in the UK have hearing impairments who may use subtitles. It also identifies 930,000 people with visual impairments and 87,000 people with both hearing and sight impairments. Those figures don't describe a marginal use case for specialist content. They describe a substantial audience that may rely on accessible alternatives to participate.


Practical rule: Treat every recording as a learning asset with multiple access routes, not as a video file with an optional caption layer.

Institutional risk matters, but it shouldn't be the only motivation. Captions help people studying in noisy environments, transcripts support revision and searching, and clear audio helps learners who depend on spoken explanations. UK guidance from Jisc on captioning and digital accessibility also reflects the reality that educational teams need to make decisions about mixed-media lectures, slides, demonstrations, and recordings already sitting in an LMS.


The sustainable approach is operational. Audit what exists, prioritise the recordings learners use most, create accessible alternatives, review them, test the player, and make the same checks part of every new upload.


Understanding Captions Transcripts and Audio Description


These three formats solve different problems. Captions synchronise text with spoken dialogue and meaningful sounds. Transcripts provide a fuller text alternative. Audio description explains important visual information that the original narration doesn't communicate.


A graphic showing three icons representing accessibility features for videos: captions, transcripts, and audio descriptions.

Captions preserve timing and context


Captions should identify speakers where that matters, include relevant non-speech audio, and appear at the right moment. A caption track that contains only dialogue can still leave out an alarm, a change in speaker, or an important audio cue. The UK guidance for creating accessible audio-visual content recommends closed captions, keyboard-accessible media players, transcripts, and audio description when visual information is essential.


Captions are also widely used beyond disability access. An Extreme Reach survey reported that 79% of UK viewers use subtitles at least sometimes, including 59% of people aged 18 to 24 who use them always or often. Those figures are reported by Advanced Television's coverage of the UK subtitle survey. In an LMS, captions support students watching without sound, reviewing unfamiliar terminology, or following a lecturer whose audio is less than perfect.


Transcripts support reading, search, and alternative access


A transcript should represent the meaningful content of the recording, including spoken words and descriptions needed to understand visuals. It can be downloadable, displayed below the player, or linked from the course page. It's particularly useful for podcast-style teaching, revision, assistive technology, and learners who need to locate a specific explanation quickly.


A transcript isn't just a caption file copied into a document. Captions are timed and segmented for viewing. A transcript can be edited into a more readable text alternative, with headings, speaker names, and descriptions of demonstrations. For teams working from online recordings, a resource such as YouTube to text transcription can help create a starting text file, but a human still needs to correct terminology and add missing context.


For a focused explanation of the caption format itself, use this guide to what closed captioning means.


Audio description handles visual dependency


Audio description narrates visually important details, such as a diagram's relationships, a lecturer's demonstration, a facial expression that changes the meaning of a statement, or the key controls used in a software walkthrough. Ofcom guidance describes audio description as commentary in the present tense that conveys information such as body language, facial expressions, and settings.


The best solution is often to write visual explanations into the original script. If a lecturer says, “The red line rises sharply,” the spoken content already carries the information a blind or low-vision learner needs. If the soundtrack leaves no space, create an alternate described version rather than forcing extra narration over speech.


Creating Accurate Captions with AI Tools


Automatic speech recognition is useful because it removes the slowest part of a large remediation project, but it doesn't remove editorial responsibility. Treat AI captions as a draft that needs review, not as a finished accessibility asset.


Start by uploading the recording to your captioning platform and generating the initial speech-to-text track. In MEDIAL, educators can use AI-assisted caption generation and edit the result in the browser. The same workflow can be applied module by module, which is more manageable than opening recordings one at a time without a prioritisation plan. For an overview of the tool category, see AI caption generation for videos.


Review the errors that matter most


Open the editor and listen while reading. Don't review the text in isolation, because a caption can look plausible while appearing too early, too late, or on the wrong speaker.


Pay close attention to:


  • Technical terminology: Compare subject-specific words with the lecture slides, glossary, or source material. Speech recognition often produces a familiar word instead of the specialist term spoken.

  • Speaker identification: Label speakers when a discussion, interview, practical session, or student contribution makes identity relevant.

  • Overlapping dialogue: Replay sections where people interrupt one another. If the system merges voices, restructure the captions so the exchange remains understandable.

  • Background noise: Check recordings with fans, room echo, keyboard sounds, or audience movement. Noise can change both word accuracy and timing.

  • Meaningful sounds: Add non-speech information when it affects comprehension, such as a warning tone, applause, or a piece of equipment starting.


UK Government Digital Service guidance says auto-generated captions should be verified, synchronised, and accurate. Don't publish unchecked output just because the player displays a CC button.


Use a repeatable quality check


Play the recording from start to finish at normal speed. Confirm that captions stay aligned, don't obscure important on-screen text, and identify speakers or sounds where needed. Then test a few difficult points again at reduced speed, especially dense explanations, equations, proper names, and sections with rapid discussion.


If your institution uses a formal accuracy threshold, document it and apply it consistently. The supplied UK guidance supports verification, but it doesn't establish a universal 95% minimum for every educational recording, so a team should define its own acceptance criteria rather than present that figure as a legal rule.


Improve the input as well as the output. Ask lecturers to use a suitable microphone, record in a quiet space, reduce echo, and speak at a steady pace. For legacy recordings, basic audio cleaning may make the draft easier to review, but it can't recover words that were never captured clearly.


Batch processing is the practical scaling mechanism. Group recordings by course or module, generate drafts together, assign review queues, and publish only after someone checks the completed files.


Deciding When to Add Transcripts and Audio Description


Not every recording needs the same treatment, but every recording needs a deliberate decision. Start with the visual-dependency test: if removing the video track makes the explanation incomplete or confusing, the production needs audio description or a spoken equivalent.


A talking-head welcome message may need captions and a transcript if the presenter communicates everything through speech. A chemistry demonstration may need captions plus descriptions of equipment, actions, measurements, and changes that aren't verbalised. A screen recording needs narration that explains both the task and the interface changes, because a transcript containing only spoken sentences may omit the actual workflow.


The UK Government accessibility explorer states that video and audio should have appropriate alternatives, including transcripts, captions for video, and audio description where narration doesn't include necessary detail. Bath's guidance on accessible video and audio distinguishes between prerecorded video with audio, prerecorded audio-only content, and video-only content, which helps teams choose the right asset rather than applying one solution everywhere.


Match the treatment to the content


Video Type

Captions

Transcript

Audio Description

WCAG Level

Talking-head lecture

Required for spoken content

Useful text alternative

Needed if essential visual information isn't spoken

WCAG 2.1 AA

Slide-based lecture

Required, including meaningful sounds

Include spoken content and visual explanations

Needed for diagrams, charts, or slide content not described aloud

WCAG 2.1 AA

Software screen recording

Required for narration and relevant sounds

Include actions and interface changes

Needed when actions aren't adequately narrated

WCAG 2.1 AA

Audio-focused lesson

Not applicable unless paired with video

Full transcript

Not applicable to audio-only content

WCAG 2.1 AA

Video without meaningful audio

Not applicable to absent audio

Descriptive transcript

Audio description or equivalent text description

WCAG 2.1 AA


A transcript can be the better revision tool for a long interview or audio lesson, while captions remain necessary for a video that contains spoken teaching. Don't force students to download a file to access information that should be available during playback.


For new recordings, build descriptions into the script. Say what a slide shows, identify the axes of a graph, describe the result of a demonstration, and avoid phrases such as “as you can see” unless the visible information is also explained aloud. This reduces later audio-description work and gives the transcript more value.


Configuring Accessible Players and LMS Integration


A perfectly edited caption file can still fail at the point of use. Students encounter the player inside a course page, discussion forum, assignment, mobile view, or embedded activity, and each context can expose different problems.


Configure the player so users can reach play, pause, volume, caption settings, transcript controls, and fullscreen mode with a keyboard. Visible focus indicators and a logical tab order matter because a caption toggle that requires precise mouse movement excludes people who interact without one. Closed captions are generally preferable to permanently burned-in text because users can turn them on or off, adjust their experience, and access the original video separately.


Test the complete route


Don't stop after checking the platform's native media page. Embed the recording in the actual Canvas, Moodle, Blackboard, or D2L Brightspace activity where students will use it. Check whether captions remain available, whether the transcript link survives the embed, and whether the player exposes controls to screen readers.


Use keyboard-only navigation first. Then test with screen readers such as NVDA and JAWS, as supported by your institutional setup. Confirm that focus moves predictably, buttons have meaningful names, and caption state changes are announced.


A numbered five-step infographic showing the process for ensuring video accessibility and caption compliance.

A common LMS failure looks harmless in an authoring view. Captions work in the original player, but disappear when the video is embedded in a discussion. Another occurs when a transcript exists as an unlinked file that students can't find or the LMS search can't index. Put the transcript beside the player with a clear label, and check the published student view rather than trusting the instructor preview.


Make publishing predictable


Use consistent metadata for course, module, topic, language, caption availability, transcript availability, and audio-description status. Store caption files with clear version names so an updated lecture doesn't accidentally retain an older track.


Where your video platform supports automation, send the caption file into a review queue after processing, then publish it to the LMS only when the reviewer marks it complete. Keep a fallback link to the source player or transcript if an embed fails. The accessibility process isn't finished until a learner can locate, operate, and understand the media in the environment where teaching happens.


Handling Live Sessions and Near-Live Recordings


Live accessibility needs its own standard. Real-time captions can help learners participate immediately, but automatic captions may struggle with specialist vocabulary, multiple speakers, poor microphones, and audience questions. Institutions should decide in advance when live automated captions are sufficient and when a professional CART service or another human-supported arrangement is necessary.


That decision should reflect the session's audience, content, interaction level, and any individual adjustment plan. A small internal briefing with one speaker and clear audio presents a different risk from a practical class with several contributors, technical terms, and safety-critical instructions. Legal minimums and learner expectations can diverge, so meeting only the minimum may still produce an unusable session.


Prepare before the event:


  • Share terminology: Give presenters and captioners names, acronyms, module terms, and slide text where the platform allows it.

  • Coach delivery: Ask speakers to use a microphone, avoid talking over one another, announce audience questions, and describe visual material aloud.

  • Prepare materials: Publish accessible slides or notes separately so participants have another route to key information.

  • Assign ownership: Name the person responsible for monitoring captions, responding to access issues, and preparing the recording afterwards.


A three-step infographic showing the process of generating accessible live captions for video content.

Near-live recordings deserve a rapid post-session workflow. If the recording will be shared with students, generate captions, review the most difficult sections, attach the transcript, and check the final player before directing learners to it. Guidance on live-streaming university lectures is useful when designing the operational hand-off between the live event and the recorded lesson.


Institutional guidance often treats live video differently from prerecorded content, but a recording that remains available creates a continuing access obligation in practical terms. Don't let “live” become a reason to leave the permanent resource inaccessible.


Your Action Plan for Accessible Video at Scale


Start with an audit, but don't try to repair the library in a random order. Create a simple inventory containing the course, recording title, audience, location, teaching period, visual dependency, caption status, transcript status, and audio-description decision.


Prioritise recordings using three questions:


  1. Usage: Which videos do students access most often?

  2. Impact: Which recordings support core learning outcomes, assessment preparation, or large cohorts?

  3. Risk: Which content is already subject to an accessibility request, audit finding, or formal institutional requirement?


Then move each priority recording through the same pipeline:


  • Audit: Decide whether captions, a transcript, audio description, or a combination is needed.

  • Generate: Create a speech-recognition draft for recordings with audio.

  • Review: Correct terminology, speakers, timing, non-speech sounds, and visual references.

  • Publish: Attach the assets in the LMS and make the alternatives easy to find.

  • Test: Use the student view, keyboard navigation, and screen-reader checks.

  • Monitor: Record unresolved issues and revisit content when the lecture changes.


Avoid promising a fixed review time per hour of footage unless your team has measured its own workflow. Audio quality, speaker count, technical vocabulary, and visual complexity change the effort substantially. Measure a representative sample, then use that internal baseline for staffing and scheduling.


For a legacy backlog, batch by module rather than by individual lecturer. Give faculty a short review checklist, provide examples of acceptable captions, and escalate only the content that needs subject expertise. Budget constraints are easier to manage when teams prioritise high-impact recordings and prevent new debt through an upload gate.


The lasting fix is simple to state: new video shouldn't enter the LMS without an accessibility decision. Make caption generation, transcript availability, visual-description checks, and player testing part of the normal publishing workflow. That turns remediation from an emergency project into routine course maintenance.



MEDIAL supports captioned uploaded and recorded media, downloadable transcripts, browser-based caption editing, and LMS-connected video workflows. To assess how it could fit your institution's production and remediation process, visit MEDIAL and arrange a trial or personalised demonstration.


 
 
 

Comments


bottom of page