SkillVaultskills Browse all 500 skills

Research · Version 1.1.0 · Reviewed 2026-08-02

A/B Test Readout Reviewer

Produce defensible evidence for validity checks and effect interpretation with evidence, explicit trade-offs, and a verification plan.

4 method steps 6 documented failure modes 5 diagnostic checks 7 quality gates

Reviews experiment results for validity before they are used to justify a decision.

₹99 one-time

Get this skill archive

What this skill helps you do

  • Validity checks
  • Effect interpretation
  • Decision framing

How A/B Test Readout Reviewer works

You provide

Suite structure, failure history, and the risk to cover

It inspects

Nondeterminism sources affecting validity checks

It decides

A effect interpretation plan at the cheapest useful level

You verify

The test fails when the behavior is broken, not only passes

What it checks first

A/B Test Readout Reviewer reviews experiment results for validity before they are used to justify a decision. Use it when the work involves Validity checks, Effect interpretation, Decision framing.

  1. Whether the test asserts behavior or implementation, because implementation-coupled tests break on safe refactors.
  2. Sources of nondeterminism: time, randomness, ordering, concurrency, network, and shared state.
  3. Whether tests share mutable state, which makes failures depend on execution order.
  4. The test pyramid balance, since a suite dominated by end-to-end tests is slow and flaky by construction.
  5. Whether a failing test failed for the intended reason, verified by making it fail deliberately.

Failure modes it recognizes

  • A flaky test caused by a fixed sleep instead of waiting for the actual condition.
  • Tests passing in isolation and failing in suite because of leaked global or database state.
  • Time-dependent assertions failing at month or year boundaries or across daylight-saving transitions.
  • Over-mocking that verifies the mock rather than the integration, so the suite passes while production breaks.
  • A test asserting on unordered collection order, which passes until the implementation changes hashing.
  • Coverage measured but assertions absent, so lines execute without being verified.

Answers it will reject

  • Retrying a flaky test to make CI green, which converts a real intermittent bug into an invisible one.
  • Chasing a coverage percentage, which produces tests that execute code without asserting behavior.
  • Writing an end-to-end test for logic that a unit test could cover deterministically and instantly.
  • Deleting a failing test to unblock a release without recording the risk that was accepted.

Decision rules it applies

  • Choose the cheapest test level that can actually observe the failure mode.
  • A flaky test is a defect in the test or the system; quarantine with an owner and a deadline, never ignore.
  • Assert on observable behavior and public contracts so refactors stay free.
  • Every bug fix gets a test that fails before the fix and passes after it.

Evidence it asks for

  • Run the suite in randomized order to expose inter-test dependencies.
  • Track flake rate per test over time rather than treating each failure as isolated.
  • Verify a new test fails when the behavior is broken, not only that it passes when correct.

The method inside

  1. Define the research question and unit of analysis
  2. Create a transparent coding or extraction framework
  3. Preserve source traceability and negative evidence
  4. Separate findings, interpretation, limitations, and applicability

Deliverables

  • Validity checks evidence table
  • Effect interpretation findings with negative cases
  • Decision framing limitations and next-research plan

Evidence requirements

  • Source documents, transcripts, data, and research question
  • Sampling method, population, and collection context
  • Known limitations, contradictory cases, and analysis criteria

Quality gates

  • Every material claim traces to supplied evidence or is labeled as a hypothesis.
  • The response follows the declared deliverable contract.
  • No execution, access, measurement, or verification is invented.
  • Secrets and personal data are redacted rather than repeated.
  • The user receives a concrete independent verification step.
  • The relevant failure modes in this domain were considered rather than only the reported symptom.
  • No listed anti-pattern was recommended as a solution.

Example task

Input

Our test shows a 3 percent lift with p under 0.05. Are we good to ship?

Expected output

Check sample ratio mismatch and pre-period balance first, because a broken split produces a significant result from nothing. Then compare the confidence interval against the effect you would act on: a 3 percent point estimate with a wide interval may not clear that bar...

Boundaries and compatibility

Ideal for

  • Validity checks: produce a decision or artifact grounded in supplied evidence.
  • Effect interpretation: produce a decision or artifact grounded in supplied evidence.
  • Decision framing: produce a decision or artifact grounded in supplied evidence.

Out of scope

  • Fabricating sources, participants, or findings
  • Claiming representativeness without a sampling basis

Agent compatibility

  • GitHub Copilot custom agents
  • Claude Agent Skills / SKILL.md
  • Any instruction-following chat model

Tool policy: Advisory by default. No tools are assumed. If the host provides tools, use read-only evidence gathering unless the user explicitly approves a scoped write or execution action.