Resonant IQ Help & Docs
Concepts QA Insights Open Resonant IQ
QA / Building scorecards & rubrics
QA workspace

Building scorecards & rubrics

The AI doesn't invent what “good” means — it scores against a rubric you write. Get the scorecard right and every score downstream inherits your definition of quality.

6 min read
For QA managers & admins
QA Manager Admin

A scorecard is your rubric, written in Resonant IQ. It's an ordered set of criteria — and each criterion has a type that decides how it's scored: some are judged by the AI, some are mechanical checks on timing, and some pull in a customer's own feedback. Together they define what quality means for your team, and every conversation is scored against exactly this set.

The rubric is yours — the scoring is automatic.
You define the criteria once; the auto-approve model applies them to every conversation without anyone pressing a button.

Three kinds of criterion

Every criterion is one of three types. Choosing the right type is most of the design work — it decides whether the AI judges the criterion, a clock does, or your customers do.

Qualitative
The AI reads the conversation against your description and returns a 0–100 score with a plain-language explanation. The only type your wording actively steers.
Threshold
A timing rule — first response, resolution, or reply time — scored pass / fail against a limit you set, and reported as a pass rate in analytics.
Customer Feedback
Pulls an imported signal like CSAT or NPS from a metadata field and normalizes it to 0–100. Counted in the weighted score only if you opt it in.

Only qualitative and threshold criteria carry weights, and those must total exactly 100%; customer-feedback criteria sit outside that total unless you include them. Any criterion can also be marked Critical — if it fails, its conversation is always flagged for a person, whatever the overall score.

Anatomy of a scorecard

A working scorecard usually mixes the types. The weighted criteria roll up into the conversation's overall score; a customer-feedback criterion can ride alongside without diluting it. Here's one mid-build.

Settings · Scorecard Reconstructed — not a live capture
Issue resolution Qualitative Critical
Did the agent actually solve the customer's problem, not just respond to it?
Weight 50%
Tone & empathy Qualitative
Did the reply meet the customer where they were, in our brand voice?
Weight 25%
First response time Threshold
First reply within 5 minutes · scored pass / fail.
Weight 25%
CSAT Customer Feedback
Imported CSAT, normalized to 0–100 · tracked alongside, not weighted.
field: csat_score
4 criteria · scored weights total 100% v3 · active
Scored weights (qualitative + threshold) total 100%; the customer-feedback criterion is tracked but left out of that total here. For qualitative criteria, the description is the instruction the AI follows — the sharper it is, the more consistent the scores.

Writing criteria that score well

This is about qualitative criteria — the ones the AI judges from your words. (Threshold and customer-feedback criteria are mechanical: you set a time limit or point at a feedback field, so they need no description.) For the qualitative ones, a few habits make a rubric that scores the way you'd score it yourself:

Describe observable behavior. “Confirmed the fix worked before closing” scores more reliably than “was helpful.”
Keep criteria independent. Overlapping criteria double-count the same thing and muddy where a low score came from.
Attach your brand voice. The brand voice guidelines you set give tone criteria something concrete to measure against.
Recalibrate against real reviews. After the first ~50 human corrections, compare AI and human scores and tighten the descriptions where they diverged.

Versioning keeps scores honest

Every score is tied to the scorecard version in effect when it was scored. Editing criteria, weights, or customer-feedback inclusion and hitting Save new version publishes an immutable version — the editor won't let you publish until scored weights total 100%. Past conversations keep the scores they were given under the old version, so history isn't silently rewritten under a new definition of quality.

That's what makes an average trustworthy: a rising score means the work improved, not that the bar moved. When you do change the rubric, expect a short recalibration window while new scores settle against the new criteria. From there, act on the scores in Working the Queue and Scores, flags & coaching.

← Previous
Scores, flags & coaching
Next up
Timeline & account health →