FairGap®

Sampling plans that boards misunderstand in bias audits

Why sampling design determines whether a bias-audit finding can be reconstructed, and the questions boards should ask before accepting a headline ratio.

Boards often treat the sampling plan in a bias audit as a technical appendix. That is a mistake. Sampling determines whose outcomes enter the measurement, which periods are treated as representative, and how much confidence anyone can place in a disparity finding. When directors ask only whether the audit "passed" and leave the sample design to the footnotes, they are asking for a conclusion without the conditions that make the conclusion usable.

This note is written for boards and counsel who review bias-audit materials for automated employment tools and adjacent decision systems. It is not legal advice. The point is narrower: several sampling choices recur across audits, and boards routinely misunderstand what those choices imply. Clearing the misunderstanding early reduces the chance that a later inquiry will expose an elegant metric sitting on a fragile base.

What a sampling plan actually decides

A sampling plan is the written record of how the audit population was constructed. It answers questions that look administrative and are, in practice, substantive. Which applicants, employees, or candidates were included? Over what date range? Were withdrawals, incomplete applications, or withdrawn offers treated as outcomes or excluded? Was the sample drawn from production traffic, a retrospective extract, a vendor test set, or some mixture? Were protected-category labels observed, imputed, self-reported, or unavailable for part of the population?

Those choices are not neutral packaging around a metric. They are the metric's boundary conditions. A selection-rate ratio computed on a clean, complete, recent cohort can look very different from the same ratio computed on a mixed historical window that includes policy changes, seasonal hiring spikes, or a vendor model version that no longer runs in production. Boards that receive a single headline ratio without the sampling frame are reviewing a number that cannot be reconstructed later.

Counsel tends to see this immediately when document requests arrive. The first useful question is rarely "what was the adverse-impact number?" It is "on whom, under which decision definition, for which period, and with which exclusions?" If the sampling plan cannot answer those questions in writing, the finding is difficult to defend as a finding.

Five misunderstandings that recur

The first misunderstanding is that larger samples are always better samples. Volume helps when the population is stable and the decision process is unchanged. Volume hurts when the larger window mixes distinct regimes: a new sourcing channel, a revised interview rubric, a model refresh, or a hiring freeze that alters who reaches the scored stage. A board that praises "n of fifty thousand" without asking whether the fifty thousand describe one process or several is mistaking bulk for coherence.

The second is that a random sample of historical records is a substitute for a representative sample of the decision that matters now. Randomness can reduce certain forms of selection bias in extraction. It does not repair a mismatch between the audited process and the live process. If the tool's features, thresholds, or human override practices changed after the sample window closed, a carefully randomized historical draw still measures yesterday. Boards should ask whether the sample corresponds to the system currently in use, not only whether the draw was random.

The third is that vendor-provided evaluation sets are interchangeable with production samples. Vendor sets can be useful for method checks and for comparing versions under controlled conditions. They are often poorly matched to the organization's applicant mix, job families, geography, and human review practices. An audit that reports strong fairness metrics on a vendor set and weak or incomplete coverage of production traffic has told the board something about a laboratory condition, not necessarily about the employment process the organization actually runs.

The fourth is that missing demographic data can be treated as a minor data-quality footnote. In many organizations, category coverage is uneven across sources, time periods, and stages of the funnel. If the audit drops records with missing labels, the remaining sample may skew toward populations more likely to self-report or more likely to appear in systems that collect those fields. If the audit imputes categories, the imputation method becomes part of the finding and must be documented with the same care as the disparity metric itself. Boards that accept "incomplete demographics" as a shrug are accepting an unexamined filter on the people whose outcomes count.

The fifth is that stratified sampling is a guarantee of fairness rather than a design choice with tradeoffs. Stratification can ensure that smaller groups are present in sufficient numbers for stable estimates. It can also overweight rare cells relative to their share of the actual decision population, which matters when the board later asks what the organization experienced in operation rather than what a balanced design could detect. Stratification is often the right tool. It is not a moral seal. The plan should state why strata were chosen, how weights were applied if they were, and what question the design was built to answer.

What boards should ask before they accept the finding

Directors do not need to become statisticians. They do need a short set of questions that force the sampling plan into the open. Ask for the population definition in one paragraph, including inclusions and exclusions. Ask for the date range and for any material process or model changes inside that range. Ask whether the sample is production, vendor, or mixed, and what share of live decisions the production portion covers. Ask how protected categories were obtained and what fraction of records lacked usable labels. Ask whether the sample size for each reported group is large enough that a modest shift in a few outcomes would not reverse the headline conclusion.

Those questions sound procedural. They are the difference between a board that can stand behind an audit summary and a board that has merely received one. In regimes that attach independent audit, notice, or publication duties to automated employment tools, including frameworks such as NYC Local Law 144, the documentary quality of the sample design is part of the compliance posture, not an academic preference. Even outside a specific statute, employment and consumer counsel tend to return to the same reconstruction problem: can a careful reader rebuild the measured population from the file?

A practical discipline helps. Treat the sampling plan as a first-class artifact in the audit package, alongside scope, metrics, independence, and remediation. Require that the plan state its purpose in plain language: detection of disparity under current operating conditions, comparison across model versions, or historical baseline for a policy change. Different purposes justify different designs. Confusion arises when a design built for one purpose is presented as answering another.

Boards often want comfort: a clean ratio, a pass or fail, a sentence that says the tool was reviewed. Sampling plans interrupt that comfort because they expose contingency. That interruption is useful. A disparity finding that cannot survive a question about who was counted is not a finding the organization can rely on when the question arrives from outside the room.

The corrective habit is simple and slightly austere. Read the sampling plan before the scorecard. Ask whether the people in the sample are the people the decision system is affecting now. Ask what was left out and why. Ask whether missing labels, vendor sets, or mixed time windows changed the story. If those answers are clear in writing, the metrics that follow have a chance of meaning what the board thinks they mean. If they are not, no amount of polish on the final ratio will repair the gap.