How call transcription works: the engine everything else stands on

    How it works6 min readPublished

    TL;DR

    • Transcription is the base of the whole chain: analysis, scoring, alerts and coaching. A weak transcript turns them all into guesswork.
    • The pipeline: audio capture from telephony, a speech model converting voice to text, speaker separation, then punctuation and timestamps.
    • Accuracy is measured as word error rate (WER) on your own real calls, not studio recordings. What matters is meaning-changing errors.
    • Below 90% accuracy, automated analysis does more harm than good by creating false confidence. Above 95% you can build QA, scoring and coaching on it.

    When evaluating a conversation intelligence system, everyone looks at the dashboards and insights. But every layer, from the score to the compliance alert, is built on one artifact: the transcript. If it is accurate, everything above works. If it is weak, everything above guesses. So it is worth understanding how it works inside.

    Step 1: from telephony to the engine

    It starts with the recording: the system receives audio from your existing telephony, usually as two separate channels, one for the rep and one for the customer. That separation matters: with each speaker on their own channel, attributing words to the right person is almost guaranteed. With a single mixed channel, the speaker-separation stage has to work much harder.

    Step 2: voice to text

    The speech model splits audio into short segments and recognizes what was said in each, based on training over enormous amounts of real speech. Two things determine quality: what speech the model was trained on (real phone calls with noise and accents, or clean recordings), and in which language. A model trained mostly on English will underperform on spoken Hebrew no matter how advanced it is. We wrote a separate guide on the Hebrew challenge.

    Step 3: diarization, punctuation and timestamps

    • Speaker separation (diarization): who said each sentence. Critical, because "I will get back to you tomorrow" is a promise from the rep and a brush-off from the customer.
    • Punctuation and sentence splitting: without it, semantic analysis receives an unstructured word stream.
    • Timestamps: every sentence links to the exact second in the recording, letting you jump from insight to moment.

    How accuracy is really measured

    The standard metric is word error rate: how many words are wrong, missing or invented relative to what was actually said. The number alone misleads, though: an error on a filler word changes nothing, while an error on a price, a product name or a negation changes everything. The right test is manual and simple: take five of your real recordings, including a noisy call and a heavy accent, read the transcript against the audio and count meaning-changing errors.

    The threshold to know

    At 85% accuracy one word in seven is wrong, and script or compliance checks become a lottery. Below 90%, analysis produces false confidence and does more harm than good. Above 95%, as a Hebrew-first system like Saleso delivers, you can build QA, scoring and coaching on the transcript.

    What an accurate transcript unlocks downstream

    Once there is accurate text with speakers and times, everything else opens up: stage and objection detection, script milestone checks, mandatory-disclosure detection, per-call scoring, automatic CRM summaries and real-time whispers. That is why, when choosing a system, the transcription question comes before any feature question: no feature is better than the transcript it stands on.

    Frequently asked questions

    What is the difference between real-time and post-call transcription?

    Post-call transcription runs on the full recording and reaches maximum accuracy. Real-time transcription must respond within seconds and is technically harder. A complete system uses both: real time for live whispers and alerts, full transcription for analysis, scoring and documentation.

    How do background noise and accents affect accuracy?

    A model trained on real phone calls handles floor noise and varied accents far better than one trained on clean recordings. Which is exactly why you must test any system on your own recordings, hard cases included, not on the vendor's demo.

    Does transcription store the recording itself?

    Recordings stay in your or your telephony provider's infrastructure, as today. The analysis layer works on them under the permissions and security rules you define, and every transcript links back to the source recording with timestamps.

    Instead of reading about it, see it on one of your own calls.