SkillVaultskills Browse all 500 skills

People Operations · Version 1.0.0 · Reviewed 2026-08-02

Structured Interview Evidence Scorer

Make competency evidence coding and panel calibration with evidence, explicit trade-offs, and a verification plan.

4 method steps 4 documented failure modes 4 diagnostic checks 7 quality gates

Scores interview notes against predefined competencies using only behavioral evidence, separates observation from inference, and highlights inconsistent or biased panel reasoning.

₹149 one-time

Get this skill archive

What this skill helps you do

  • Competency evidence coding
  • Panel calibration
  • Bias and inconsistency detection

How Structured Interview Evidence Scorer works

You provide

Interview notes, rubric, or survey data with response rates

It inspects

Evidence specificity and rubric consistency for competency evidence coding

It decides

A panel calibration judgment with gaps marked rather than guessed

You verify

Score distributions compared across evaluators for drift

What it checks first

Structured Interview Evidence Scorer scores interview notes against predefined competencies using only behavioral evidence, separates observation from inference, and highlights inconsistent or biased panel reasoning. Use it when the work involves Competency evidence coding, Panel calibration, Bias and inconsistency detection.

  1. Whether evaluation evidence is behavioral and specific or impressionistic.
  2. Whether the same standard was applied across candidates or drifted between them.
  3. Sample size and anonymity conditions behind any survey conclusion.
  4. Whether a theme reflects a widespread issue or a small vocal group.

Failure modes it recognizes

  • Scores assigned before evidence is recorded, so the evidence is written to justify the score.
  • Free-text survey themes dominated by the most articulate respondents rather than the most common view.
  • Comparison across interviewers who applied different implicit bars.
  • Anonymity promised but breakable through small-group segmentation.

Answers it will reject

  • Reporting sentiment percentages from a low-response survey as if representative.
  • Using a rubric as decoration while the decision is made on overall impression.
  • Aggregating feedback in a way that identifies individuals in small teams.

Decision rules it applies

  • Record the behavioral evidence before assigning any score.
  • Apply one rubric consistently and flag where evidence is insufficient rather than guessing.
  • Protect anonymity by suppressing segments below a minimum response threshold.

Evidence it asks for

  • Compare score distributions across interviewers to detect a drifting bar.
  • Report response rate and denominator alongside every survey finding.
  • Attach representative verbatims to each theme without identifying detail.

The method inside

  1. Define the decision criterion before reading the evidence
  2. Separate observation from interpretation and bias
  3. Check consistency across people, segments, or reviewers
  4. Produce actionable language while preserving confidentiality

Deliverables

  • Competency evidence coding evidence assessment
  • Panel calibration consistency findings
  • Bias and inconsistency detection action-ready revision

Evidence requirements

  • Role rubric, policy, survey, or review artifact
  • Observable behavior and outcomes
  • Relevant context with unnecessary personal identifiers removed

Quality gates

  • Every material claim traces to supplied evidence or is labeled as a hypothesis.
  • The response follows the declared deliverable contract.
  • No execution, access, measurement, or verification is invented.
  • Secrets and personal data are redacted rather than repeated.
  • The user receives a concrete independent verification step.
  • The relevant failure modes in this domain were considered rather than only the reported symptom.
  • No listed anti-pattern was recommended as a solution.

Example task

Input

Score these panel notes against our leadership rubric and flag ratings that are not supported by interview evidence.

Expected output

Execution has three concrete high-scope examples; strategic thinking is rated “strong” but supported only by hypothetical answers. Two interviewers penalized communication style without tying it to the rubric...

Boundaries and compatibility

Ideal for

  • Competency evidence coding: produce a decision or artifact grounded in supplied evidence.
  • Panel calibration: produce a decision or artifact grounded in supplied evidence.
  • Bias and inconsistency detection: produce a decision or artifact grounded in supplied evidence.

Out of scope

  • Making employment decisions without accountable human review
  • Inferring protected characteristics or psychological states

Agent compatibility

  • GitHub Copilot custom agents
  • Claude Agent Skills / SKILL.md
  • Any instruction-following chat model

Tool policy: Advisory by default. No tools are assumed. If the host provides tools, use read-only evidence gathering unless the user explicitly approves a scoped write or execution action.