Skills Quality assurance
Gold set management
A gold set is the reference standard for whether graders apply the rubric the same way. Most programmes build one once, reuse it until everyone has memorised the verdicts, and then report calibration numbers that measure recall, not agreement.
npx rulebase-skills install cx-gold-set-managementWhen to use it
Reach for this when someone says any of these — they are the phrases the skill itself triggers on:
- “refresh our calibration set”
- “graders know the gold set by heart”
- “how often should we rotate calibration cases”
- “our calibration scores look perfect but production agreement is bad”
How it works
The method, in the order the skill runs it. The full procedure — tables, worked examples and the edge cases — is in the skill itself.
What a gold set is for — and what it is not
For: measuring agreement on criteria, onboarding graders, validating rubric changes, comparing grader or model versions on a fixed reference.
Step 1: Size for re-scoring, not coverage
The gold set must be small enough that 3+ graders can independently score the full set in one session, and that you can rotate it without a quarter-long project.
Step 2: Build with stratified sampling
A gold set must vary on the dimensions where graders disagree in production. A set drawn only from "typical" email tickets will miss voice, edge cases, and the criteria that actually break.
Step 3: Annotate as a specification, not a score sheet
Each gold item needs: Reference verdicts per criterion, with evidence quotes, Notes on ambiguity where reasonable graders could disagree.
Step 4: Set a rotation cadence before memorisation
Rotation is not optional. Define cadence by exposure, not calendar alone.
Step 5: Retire cases that stop working
Remove an item when: The conversation no longer reflects current policy or product, PII or customer context makes it unusable.
Step 6: Run calibration sessions correctly
Independence rules (non-negotiable): Graders score before seeing reference verdicts or each other's work, No "walkthrough" of the gold set in the same week as measurement, Sequential review is training, not calibration.
Related skills
Free and open source, and vendor-neutral — it reads the conversations from whichever helpdesk you already run. Browse all 149 skills · connect your helpdesk over MCP · source on GitHub
