AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Performance Calibration Pack

skill-we-are-move-claude-skills-for-ta-and-people-teams-performance-calibration-pack · by we-are-move

Designs and prepares a performance calibration cycle end to end — session structure and grouping, the manager pre-work brief, a full facilitator run sheet with the questions that move a discussion from impression to evidence, in-the-room bias interrupters, the distribution and consistency checks to run before and after, and the follow-through and appeals route. Use when someone says they are "run…

No reviews yet
0 installs
0 views
view→install

Install

$ agentstack add skill-we-are-move-claude-skills-for-ta-and-people-teams-performance-calibration-pack

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-we-are-move-claude-skills-for-ta-and-people-teams-performance-calibration-pack)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
3d ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Performance Calibration Pack? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Performance Calibration Pack

Produces the pack a People team runs a calibration cycle from: session design, manager pre-work brief, facilitator run sheet, bias interrupters, consistency checks, and the follow-through. The test of the output is that someone who has never facilitated a calibration session can run one on Tuesday and produce ratings that mean the same thing across every manager in the room.

Calibration is the mechanism that makes a rating scale mean anything. Run well it is the highest-integrity moment in the people calendar — the one point where a company checks its own judgement against evidence. Run badly it is a room where the most senior or most articulate manager's ratings survive intact and everyone else's get moved to fit, which is worse than no calibration at all: it launders inconsistency as rigour and gives the resulting pay and promotion decisions a legitimacy they have not earned.

What you need to start

Minimum viable input: the rating scale in use, roughly how many people are being rated, and when the sessions are. That is enough to design the whole thing.

If the user cannot answer even that, do not stall. Assume a four-point scale, assume ratings feed pay, design the pack, and label the assumptions at the top. A pack built on stated assumptions that the user corrects in five minutes beats a perfect pack they never got to because the conversation turned into a form.

Ask in small batches, two or three at a time, reflecting back what you have before asking more. Where the environment supports structured multiple-choice questions, use them for the scale, the consequences and the session level — those are picks, not essays.

Batch one — the scale and its consequences. What is the rating scale and do the points have written definitions? Do ratings drive pay, promotion, both, or neither? This is the first question because it sets the stakes of everything that follows: a rating that only informs a development conversation can tolerate looseness, a rating that sets a bonus multiplier cannot.

Batch two — the population and the calendar. How many people are rated, across how many managers, and how are they distributed by level and function? When do managers submit and when do the sessions run? Population size drives session count and time per person; the calendar drives whether the pre-work deadline is real or aspirational.

Batch three — the room. How many sessions, at what level do they run (team, function, company), and who chairs them? Has this population been calibrated before, and what went wrong? The failure story is the most useful answer in the whole intake — it usually names the exact intervention the run sheet needs to carry.

Useful but never required: the current rating distribution, last cycle's ratings by manager, the career framework or level definitions, the promotion criteria, and the pay matrix.

Process

1. Fix the purpose of the session before designing it

Say plainly what calibration is for, because most rooms drift without it: to check that the same evidence produces the same rating in different managers' hands. It is not a meeting to decide ratings from scratch, not a talent review, not a succession discussion, and not a budget meeting. Each of those is legitimate and each destroys calibration when merged into it — the moment the room starts allocating a bonus pool, the discussion stops being about evidence and becomes about who gets what. Schedule those separately, afterwards, using the calibrated ratings as input.

2. Design the room

Who is in it. The managers whose people are being discussed, a facilitator who owns no ratings in the session, and a People partner who holds the data and the checks. Nobody else — every additional observer changes what managers are willing to say about their own people, and honesty about a weak performer is the thing calibration most needs.

Who should not be. The rated person's skip-level, where their presence would silence the direct manager. Anyone attending to protect a favourite. Executives dropping in for one name — a senior leader who wants to intervene on one person does it through the facilitator beforehand, on the record, and the room evaluates it like any other input.

The facilitator owns no ratings in the room they run, because a facilitator defending their own team cannot credibly challenge anyone else's. Where a company is too small for that, calibrate the facilitator's own team in a different session chaired by someone else.

3. Group the population by level and function, never by manager

Group people who are doing comparable work at a comparable level. Calibration is a comparison exercise and the comparison is only meaningful between like and like — a senior engineer and a junior marketer share a rating scale but nothing else.

Manager-by-manager walkthroughs are the default in most tools and they defeat the purpose. Each manager presents their own internally ordered list, so the room evaluates that manager's consistency with themselves — the one thing that was never in doubt — while the inconsistency between managers, which is what you convened to find, stays invisible. Interleaving by level forces the actual question: is this manager's "exceeds" the same as that manager's "exceeds"?

Practical grouping rules: one session per level band per function where the population supports it, and where it does not, group adjacent levels while holding the level distinction explicitly in the discussion. Cap a session at what fits the time — two focused sessions beat one that runs long and starts nodding people through at the ninety-minute mark. Where someone spans functions or changed manager mid-cycle, place them where the work was done and brief both managers to attend that session.

4. Set the time budget honestly, and decide what is discussed

Work from the total time available, not from an aspiration: budget a few minutes per person on average, and spend it unevenly on purpose. Not everyone needs discussion, so sort before the session.

  • Discuss — the top and bottom of the scale, every rating change proposed since the

last cycle, anyone whose rating carries a promotion or a performance-management consequence, anyone whose manager is new, anyone flagged by the pre-session checks.

  • Confirm quickly — mid-scale ratings with a clear rationale and no flags. Read the

name, state the rating, pause for challenge, move on. Ten seconds each is honest.

  • Never rubber-stamp silently. Say aloud that a name is being confirmed without

discussion so anyone can stop it. An unspoken name is an unchallenged name.

Tell the room the budget at the start and hold it. Sessions that overrun do not distribute the shortfall evenly — they discuss the first third properly and rush the rest, which means rating quality ends up depending on the running order.

5. Settle the distribution question — with the argument, not an assertion

This is the most contested design decision and the user will be challenged on it, so give them the reasoning rather than a position to defend without one.

The case for distribution guidance. Left alone, ratings inflate. Managers rate generously because it is the path of least conflict, because a high rating is a cheap way to reward someone when pay is constrained, and because nobody wants to be the manager whose team scored lowest. Once most of the population sits in the top two points, the scale stops carrying information: pay cannot be differentiated, promotion signals nothing, and the genuinely exceptional are indistinguishable from the merely fine. Guidance counteracts that drift and gives managers cover — "the expectation is that most people are performing well, and the top rating is rare" is a much easier conversation than a manager alone deciding to be the strict one.

The case against hard quotas. A quota applied to a small population produces injustice mechanically. A genuinely strong team of six does not contain a low performer, and forcing one out of it means telling someone their rating reflects their team's size rather than their work. Managers work this out quickly, and the rational response is to game it — importing a weak hire to absorb the low slot, trading names across teams, rating strategically for next cycle. Quotas also collide badly with small samples: the smaller the group, the more its distribution is noise. And the reputational cost is durable — a quota teaches people that their rating is a function of the distribution rather than of their work, and once that reading takes hold they discount every part of the performance process, including the parts that were sound.

The recommendation: guidance, not quota, with a requirement to justify departures. Publish an expected shape for the population as a whole, state that it describes the company and not any individual team, and require any manager whose team departs materially from it to explain why in the session. The justification requirement is what makes guidance bite — without it guidance is a suggestion; with it, a manager who rates six of eight people at the top must produce six evidenced cases, which either holds up or collapses under one round of questions. Two implementation notes:

  • Apply the shape at the level the sample is meaningful — usually function or company,

not team. A shape enforced on a team of six is a quota by another name.

  • Never move an individual's rating to fix a shape. The shape is a diagnostic that says

"look here". The only legitimate reason to change a rating is that the evidence does not support it. If a manager's distribution looks wrong and every individual case holds up under questioning, the distribution was right and the guidance was wrong for that team.

6. Build the manager pre-work brief

The session succeeds or fails on what managers bring. Specify it precisely and give it a deadline that is genuinely before the session, not the night before.

What each manager brings for each person: the proposed rating, a written rationale referencing the framework or level expectations, two or three specific pieces of evidence from across the whole period, and their view on the two questions the room will ask — compared to who, and what would have made this a higher rating.

The rating rationale standard, which is the part most managers get wrong:

  • Evidence, not adjectives. "Consistently excellent" is not a rationale. "Led the

billing migration, unblocked it when the vendor slipped, delivered three weeks late against a plan that had assumed no slippage" is.

  • Referenced to the level, not the person's own history. The question is whether they

met the expectations of the level they are at. Improvement is a development conversation; the rating is a standard.

  • Spanning the whole period. Require at least one substantial piece of evidence from the

first half. This single requirement does more against recency bias than any amount of in-room facilitation.

  • Written as if someone else will read it. They may — an appeal, a future manager, a

court. Professional, factual, about work.

Read references/bias-interrupters.md for the language patterns to warn managers about before they write, particularly how personality-based and achievement-based language gets distributed unevenly across a population.

The deadline discipline: submissions close far enough ahead that the People partner can run the checks and the facilitator can build the agenda from them. State the consequence plainly — a person whose rationale is not submitted is discussed last, from whatever their manager can say live, and the facilitator records it. That is not a punishment, it is what is physically possible.

7. Run the pre-session consistency checks

Cut the proposed ratings before the session and bring the cuts into the room. Their purpose is to build the agenda: they tell the facilitator which names to spend time on. Cut by manager, level, function, tenure band, full-time versus part-time and any reduced or flexible arrangement, people who had leave during the period, people who changed manager mid-cycle, and new hires who joined part-way through. Then, separately and carefully, by any demographic dimension the organisation lawfully monitors.

How to frame the demographic cut, and this matters. Its purpose is to surface a pattern for investigation — a group rated systematically lower is a signal that something in the process is not working, and the investigation looks at the evidence, the rationales and the managers, not at the individuals. It is never a basis for changing an individual's rating. Adjusting any individual's rating because of their sex, race, age, disability, or any other protected characteristic is unlawful discrimination, regardless of the direction of the adjustment or the good intention behind it. Say this in the pack, in those terms, so that nobody in the room can misread the chart as an instruction to rebalance.

The legitimate responses are: re-examine the rationales in the affected group against the standard, check whether the work allocation that preceded the ratings was equitable, and identify whether particular managers account for the pattern. Where a pattern is material or recurring, run the analysis through legal counsel — in several jurisdictions analysis conducted at counsel's direction attracts privilege, and analysis run casually in a spreadsheet does not. references/bias-interrupters.md gives the specific cut for each bias pattern and what a concerning result looks like.

8. Prepare the facilitator run sheet

Read references/facilitator-script.md in full before producing this section. It carries the opening framing, the per-person protocol, the evidence questions, and the handling for the moments that decide whether a session is worth anything: the manager who cannot evidence a rating, the dominant voice, the trade, the reopened decision, and the room that has gone quiet because it learned that challenge is expensive. The run sheet in the pack should be usable standing up — timings, opening words, the question set per person, the interventions in escalation order, and the closing.

9. Design the follow-through before the session, not after

The session is the middle of the process, not the end. Specify:

  • Who tells whom, and when. The rating is delivered by the person's own manager, in a

conversation, before it appears in any system. A rating that arrives by notification before a manager has explained it is the single most reliable way to turn a fair rating into a grievance.

  • What managers may say about calibration. Give them the line: the rating was reviewed

with other managers to make sure the standard is applied consistently across the company. Managers may not attribute a rating to the room ("I wanted to give you a 4 but they knocked it down") — it abandons their own accountability and tells the person their manager does not stand behind the decision. A manager who cannot defend a rating is a signal the session did not finish its job. They may not disclose other people's ratings or evidence, who said what, or their team's distribution.

  • People whose rating moved in the room. Every change gets a written reason recorded

against it, and the manager is briefed on how to deliver it beforehand. A manager who first learns the rating changed when they open the system will communicate it badly, and will be right to be annoyed.

  • The appeals route. Who hears an appeal, on what grounds, in what window, and what

outcomes are possible. Ground appeals in process and evidence — material evidence not considered, or the process not followed — rather than disagreement with the judgement, or every appeal becomes a re-litigation of the rating. The appeal is heard by someone who was not the deciding manage

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.