Not a generic scorecard, and not a sentiment number. Rubricon Score reads the whole call, marks each line of your existing rubric, and quotes the moment in the transcript it marked on — so a score can be argued with.
Manual QA is bounded by scorer hours, so it samples. A sample tells you the average and hides the individual call. Reading every call changes what the output is for.
The compliance miss, the wrong policy quoted, the abusive caller — these are exactly the calls a random sample is least likely to contain, and the ones that cost the most when they are missed.
An agent's whole month is scored, so the note is "you skipped the disclosure on nine of these calls, here they are" rather than a figure derived from four calls.
Your QA team stops marking routine calls and spends the hours on calibration and on the disputes — the work that actually needs a person's judgement.
Before any score is shown to a supervisor, the model has to agree with your own scorers on calls they marked by hand. That set stays in place afterwards as the thing new versions are re-tested against.
Around 200 past calls, marked by your scorers on your rubric, covering every language the queue takes. Their judgement is the target; the model is never the reference.
Per rubric line and per language, never pooled. A rubric can agree at 90% in English and 60% in Tamil, and a single number hides exactly that.
Every mismatch is looked at. Some are model error; some show the rubric line was ambiguous and needed rewriting anyway.
A fresh hand-scored batch each month, checked against the live scores, so drift surfaces as a number instead of as a complaint.
This is the part that is hard to assemble from two vendors. When the agent and the human team are measured on one versioned rubric, the comparison is real: you can see which call types the agent handles at or above your team's standard, and hand it those and no others.
It also means Rubricon is graded by its own audit layer, on your definition of a good call. If the agent regresses after a change, it shows up in your dashboard on your rubric — not in a status page of ours.
Nothing has to change in the call flow to find out whether the scoring is any good. That is deliberately the cheapest thing to test first.
No integration, no change to your telephony, and you keep the comparison whatever you decide afterwards.