Skills Quality assurance

Designing a QA scorecard that means something

A scorecard is a measuring instrument. Most are built as checklists of things that sound like quality, then discovered — usually a year in — to predict nothing and change no behaviour. This is the procedure for building one that does, and for diagnosing one that doesn't.

Quality assuranceScorecards and calibrationPlaybookAny helpdeskRead-only
Installnpx rulebase-skills install cx-qa-scorecard-design

When to use it

Reach for this when someone says any of these — they are the phrases the skill itself triggers on:

  • build a QA scorecard
  • our QA scores don't mean anything
  • design a quality rubric
  • everyone scores 98%

How it works

The method, in the order the skill runs it. The full procedure — tables, worked examples and the edge cases — is in the skill itself.

  1. Diagnose before you build

    If a scorecard already exists, find which of these it has. Each has a different fix, and rewriting criteria won't fix problems 1 or 5.

  2. Step 1: Name the decision the score must support

    Write it down before touching criteria. Scorecards commonly serve one of: Coaching, Assurance, Gating, Automation quality.

  3. Step 2: Choose the outcomes you want to move

    Pick one or two measurable outcomes the score should predict: CSAT / DSAT rate, Repeat contact rate within 7 days on the same intent, Reopen or escalation rate, Compliance breach count, Handle time, only where it is not in tension with the above.

  4. Step 3: Generate criteria from evidence, not from a template

    Vendor templates produce generic criteria that describe support in the abstract. Instead, read real conversations, roughly 20–30 of each: DSAT / low-rated conversations, Conversations followed by a repeat contact on the same intent, Conversations that escalated or reopened, A control set of high-rated conversations.

  5. Step 4: Filter every candidate through four tests

    Drop or rewrite anything that fails one.

  6. Step 5: Separate auto-fail from scored

    These are different mechanisms and must not share a scale.

  7. Step 6: Weight by consequence, not frequency

    Weight what is costly when absent, not what appears most often. A resolution accuracy criterion that fires on 8% of conversations but drives repeat contacts deserves more weight than a greeting criterion that applies to all of them.

  8. Step 7: Write each criterion as a decision rule

    Every criterion needs a name, the verdict options, the evidence required, and at least one pass and one fail example drawn from real conversations.

Related skills

Free and open source, and vendor-neutral — it reads the conversations from whichever helpdesk you already run. Browse all 149 skills · connect your helpdesk over MCP · source on GitHub

Review every conversation. Act on what it finds.

AI for customer operations, built for financial services. Specialist agents chase every issue to resolution and every stalled customer to activation.

Rulebase dashboard