Skills Quality assurance

Gold set management

A gold set is the reference standard for whether graders apply the rubric the same way. Most programmes build one once, reuse it until everyone has memorised the verdicts, and then report calibration numbers that measure recall, not agreement.

Quality assuranceScorecards and calibrationPlaybookAny helpdeskRead-only
Installnpx rulebase-skills install cx-gold-set-management

When to use it

Reach for this when someone says any of these — they are the phrases the skill itself triggers on:

  • refresh our calibration set
  • graders know the gold set by heart
  • how often should we rotate calibration cases
  • our calibration scores look perfect but production agreement is bad

How it works

The method, in the order the skill runs it. The full procedure — tables, worked examples and the edge cases — is in the skill itself.

  1. What a gold set is for — and what it is not

    For: measuring agreement on criteria, onboarding graders, validating rubric changes, comparing grader or model versions on a fixed reference.

  2. Step 1: Size for re-scoring, not coverage

    The gold set must be small enough that 3+ graders can independently score the full set in one session, and that you can rotate it without a quarter-long project.

  3. Step 2: Build with stratified sampling

    A gold set must vary on the dimensions where graders disagree in production. A set drawn only from "typical" email tickets will miss voice, edge cases, and the criteria that actually break.

  4. Step 3: Annotate as a specification, not a score sheet

    Each gold item needs: Reference verdicts per criterion, with evidence quotes, Notes on ambiguity where reasonable graders could disagree.

  5. Step 4: Set a rotation cadence before memorisation

    Rotation is not optional. Define cadence by exposure, not calendar alone.

  6. Step 5: Retire cases that stop working

    Remove an item when: The conversation no longer reflects current policy or product, PII or customer context makes it unusable.

  7. Step 6: Run calibration sessions correctly

    Independence rules (non-negotiable): Graders score before seeing reference verdicts or each other's work, No "walkthrough" of the gold set in the same week as measurement, Sequential review is training, not calibration.

Related skills

Free and open source, and vendor-neutral — it reads the conversations from whichever helpdesk you already run. Browse all 149 skills · connect your helpdesk over MCP · source on GitHub

Review every conversation. Act on what it finds.

AI for customer operations, built for financial services. Specialist agents chase every issue to resolution and every stalled customer to activation.

Rulebase dashboard