TL;DR
- A call score is compression: dozens of checks on one call, folded into a number you can track over time.
- A reliable score starts from your criteria: your script's milestones, objections and regulatory duties, each with a weight. Not a generic template.
- Checking is semantic, not keyword search: the system recognizes a price presentation even when phrased five different ways.
- Validate the score against human judgment on a sample, and use it as a trend. A single score is a data point; a two-week trend is a signal for action.
"The call scored 81." What does that actually mean? A call score can be the strongest management tool on the floor or a meaningless number, and the difference is not the algorithm but how it was built. Here is the process, step by step, including the most important question: how you know you can trust it.
Step 1: define what a good call is for you
There is no universal definition of a good call. On an insurance floor, an excellent call missing the disclosure sentence is a failed call; on an appointment-setting floor, closing a date inside the call outweighs everything else. So a reliable score starts with a short definition workshop: take the script, the common objections and the regulatory duties, and define a weighted criteria list. With Saleso this is part of implementation, and the criteria are yours, not a template.
Step 2: the system checks every call, semantically
For every transcribed call, the system goes criterion by criterion and asks: did this happen? The check is semantic, not literal: "it will cost you 300 a month", "we are talking 300 shekels monthly" and "this plan is 300" all count as presenting the price. That is the essential difference from old keyword-based systems, which failed exactly here: the rep rephrased, the check missed.
Step 3: weighting into one number, with a full breakdown
Criteria are weighted by the importance you defined, and every call gets an overall score alongside a breakdown: what was done, what was skipped, and at which second of the recording. The breakdown is the heart: 81 alone tells a rep nothing, but "81, because discovery was skipped" is already a one-minute lesson. A number without a breakdown is a black box, and black boxes create neither trust nor change.
How you verify the score is reliable
- Calibration against humans: early on, a reviewer scores a sample of calls in parallel with the system and you compare. Consistent gaps tune the definitions.
- Consistency: the same call must get the same score every time. This is the built-in advantage over human reviewers, who score differently on a busy day.
- Transparency: every score component can be shown with the moment in the call it rests on. If a score cannot be explained, it cannot be trusted.
- Maintenance: when the script or regulation changes, the criteria update and the score stays relevant.
What to do with the score daily
Three core uses: per-rep trends (a score declining two weeks straight is an orange flag long before conversion drops), fair rep comparison on identical criteria, and automatic alerts when a score collapses or a critical component, like a disclosure, is missing. And what not to do: punish a single score. One weak call is noise; a pattern is a signal.
The rule that matters
A call score is a coaching tool, not a punishment tool. When the score serves improvement, reps embrace it; when it becomes a stick, they start gaming it. The culture around the number matters as much as its accuracy.
Frequently asked questions
Can reps game the score?
Very hard. The check is semantic, so magic words without substance do not count, and the breakdown links to real moments in the recording any manager can open and hear. In practice, the easiest way to raise your score is simply to run a good call.
Generic or custom scoring model?
Custom, almost always. A generic score (talk ratio, keywords, sentiment) is interesting for two weeks; a score on your criteria reflects what actually moves your business, so it stays in use. It is one of the questions worth asking every vendor before choosing.
How many criteria should we define?
Start with 5 to 10 covering core stages, key objections and regulatory duties. Fewer and the score is shallow; many more and it diffuses. You can always expand after a month once you see what works.