Transcripts for audio
We create text transcripts of audio — podcasts, interviews, voice recordings: the full text of what is said. For deaf and hard-of-hearing users, and also for search engines and those who prefer to read. So that audio content is accessible and findable, rather than 'locked' in sound. Honestly upfront: a quality transcript is human transcription with checking; auto-recognition is only a draft with errors on terms, names and with noise. The cost depends on duration, new recordings require new transcripts.
Transcripts for audio — overview

Transcripts for audio is the creation of an accurate text version of audio content (podcasts, interviews, voice messages, voice lectures): the full text of speech with correct punctuation, paragraph breaks and, if needed, speaker identification. A transcript makes audio accessible to deaf and hard-of-hearing users (a WCAG criterion for audio), convenient for those who prefer to read or quickly skim, and indexable by search engines (text vs sound). Honestly about quality, this is key: a good transcript is HUMAN transcription with checking. Auto-recognition is only a DRAFT: it errs on terms, proper names, accents, with noise and overlapping voices. Publishing it as 'accessibility' without proofreading is dishonest. We can take auto as a base, but finish it manually. Honestly about volume: the price depends on the DURATION of the recording (minutes) and complexity (number of speakers, terminology, sound quality). Honestly about the process: new recordings require new transcripts — this is maintenance. Honestly about the boundary: a transcript is the TEXT of audio; for VIDEO synchronized captions are needed (811), because there the binding to time and picture matters; a transcript is a whole text. Honestly about the effect: accessibility + an SEO bonus + reading convenience; we do not guarantee direct sales growth. Honestly about access: source audio recordings are needed. An important boundary: this is audio transcripts; video captions are 811, audio description of visuals is 814. Picture this: instead of 'a podcast is inaccessible to deaf users and invisible to search' — a full text, accessible and indexable. The base price starts from 10,000 ₽ (depends on duration).
Problems we solve
- Deaf and hard-of-hearing users have no access to audio content.
- Podcasts/interviews are invisible to search engines (sound only).
- Auto-recognition produces text with errors on terms/names.
- No text version for those who prefer to read.
What's included in the Transcripts for audio service
- Accurate transcription of speech with punctuation and paragraphs
- Speaker identification if needed
- Finishing auto-recognition manually (if there is a base)
- Correct terminology and proper names
- A format convenient for publication and indexing
- Indicating boundaries (auto — only a draft; work by volume)
- Checking transcription quality
- Handover and review with you
What you get
- Audio is accessible to deaf and hard-of-hearing users
- Content is indexed by the search engine (text vs sound)
- There is a convenient text for reading and skimming
- The transcript criterion is closed (new recordings — maintenance)
How the work goes: steps
- We get the audio, clarify complexity; estimate by duration
- We transcribe, finish manually, verify terms/names
- We hand over in a convenient format, review boundaries with you
Why PDV Expert
- Fixed price and timeline — no surprises on the invoice.
- Report and recommendations in plain language — clear without a technical background.
- In touch at every step and answering questions about the result.
FAQ
Can we just run it through auto-recognition?
As a base — yes, but it is only a draft: auto errs on terms, names, accents and with noise. Publishing it as 'accessibility' is dishonest — a deaf user gets distorted text, and the search engine indexes the errors. We finish the transcription manually with checking, and that is real accessibility.
How does a transcript differ from captions?
A transcript is a whole TEXT of audio (podcast, interview) read separately from the recording. Captions (811) are time-synchronized text for VIDEO, bound to the picture. Audio needs a transcript, video needs captions. We will honestly say which fits your content.
Why does the price depend on duration?
Because transcription is work by volume: an hour of recording requires substantially more work than ten minutes, especially with multiple speakers, terminology or poor sound. We will calculate honestly by duration and complexity in advance, without surprises.
About the provider
The «Transcripts for audio» service is provided by PDV Expert — a team specialising in «Site quality». We work under contract and deliver a written report with recommendations.