FairGap®

What evidence a bias audit should leave behind for counsel

The documentary trail a careful bias audit should leave so counsel can reconstruct scope, measurement, independence, and remediation when questions arrive later.

A bias audit that ends when the PDF is emailed is unfinished work. Counsel's job, when an inquiry, dispute, or board question arrives, is not to admire the conclusions. It is to reconstruct what was measured, under what assumptions, by whom, on which data, and what the organization did with the result. If that reconstruction depends on memory and Slack threads, the audit's practical value collapses at the moment it is most needed.

This note describes the evidence trail a careful bias audit should leave for senior compliance counsel. It is written from audit practice, not as legal advice. Jurisdictions differ, and counsel will decide what retention, privilege, and disclosure posture fits the organization. The underlying point is narrower: an audit that cannot be defended as a process is difficult to defend as a finding.

The file counsel will actually open

When counsel opens the audit file months later, they are rarely looking first for a colorful dashboard. They are looking for a chain that answers a hostile timeline. What system was in scope? What decision did it make? What population was measured? Which categories were available and which were missing? Which metrics were computed, with which formulas? Who performed the work, and were they independent of the vendor and the development team where independence was required? What version of the tool was live when the data was pulled? What changed between the audit window and today?

Those questions sound administrative. They are the difference between a finding that holds and a finding that dissolves under document requests. Regimes such as NYC Local Law 144 make the point concrete: covered use of an automated employment decision tool carries independent audit, publication, and notice obligations that are themselves documentary. Even outside that statute, employment, credit, and consumer protection counsel tend to ask for the same underlying materials because the theories of harm turn on population, decision, and measured effect.

Scope artifacts before any metric

The first layer of evidence is the scoping record. It should state, in writing, the system identity (including vendor name and internal system ID if both exist), the decision the system produces or substantially shapes, the business process in which that decision sits, and the jurisdictions believed to apply. It should also record what was excluded from scope and why. Exclusions that exist only in a kickoff deck tend to be forgotten; exclusions that exist in a signed scope memo can be explained.

The population definition belongs in the same layer. Counsel needs to know whether the audit used historical production outcomes, a constructed test set, or a hybrid, and how the auditors judged representativeness. They also need to know the observation window, any filters applied (for example, removing withdrawn applications), and whether downstream human overrides were treated as part of the decision or as a separate step. Ambiguity here is how two competent readers produce incompatible impact ratios from "the same" audit.

A related artifact is the data dictionary for the extract: field definitions, protected-category coding rules, handling of missing or declined self-identification, and any joins across ATS, HRIS, or vendor logs. Missingness is not a footnote. In many organizations, category coverage is incomplete, and the audit's limitations section is only credible if the underlying coverage rates are preserved.

Measurement artifacts that survive turnover

The second layer is the measurement package. At minimum, counsel should be able to recover the exact metric definitions used (selection rate, scoring rate, impact ratio, or other required measures), the category cuts performed, the cell counts behind each rate, and the computed ratios. Rounded presentation tables are not enough. The underlying counts matter when someone later challenges a denominator.

The package should also preserve the code or query logic that produced the tables, or an equivalent reproducible procedure, along with the input extract identifiers (file hashes or warehouse snapshot IDs). "We ran it in a notebook" without a pinned version is how audits become non-reproducible within a quarter. Reproducibility is not an academic preference here. It is how the organization shows that the published or reported numbers were not improvised.

Where regulations prescribe a publication format, the file should contain the published artifact itself, the date it was posted, the URL, and a record of updates. Where notice to candidates or consumers is required, the file should contain the notice text, the channel, and evidence of timing relative to use of the tool. Those items are easy to treat as "compliance ops" rather than "audit." Counsel experience suggests they are examined together.

Independence, versions, and change control

The third layer concerns who did the work and what system they audited. Independence requirements vary by regime, but the evidence pattern is consistent. Retain the engagement letter, conflict disclosures, and a short statement of the auditor's relationship (or lack of relationship) to the vendor and the development team. If the organization relied on a vendor-sponsored study, label it as such and do not file it as if it were the organization's independent audit.

Version control belongs beside independence. Models, prompts, thresholds, and workflow rules change. An audit of version 3.2 does not automatically speak to version 4.1. The file should identify the production version or configuration hash in scope, the date range of data, and any material changes known at the time. A lightweight change log that product and ML already maintain is often sufficient if it is linked into the audit record. What fails is the oral assurance that "nothing important changed."

Findings, response, and the path after the report

The fourth layer is the organizational response. Counsel will want the findings in their original form, not only the executive summary. They will want a written management response: which findings are accepted, which are disputed, which are remediated, which are accepted as residual risk, and who owns each item. They will also want dated evidence of remediation where remediation is claimed, such as a configuration change ticket, a threshold update, or a process redesign with an effective date.

This layer is where many programs are thin. The audit report is archived. The remediation lives in a product backlog with no cross-reference. Months later, nobody can show whether the impact ratio that concerned the committee was addressed in production. A simple findings register that ties each finding to evidence of closure is usually enough. Complexity is optional. Traceability is not.

Privilege and work-product questions should be handled by counsel, not by the audit team improvising labels. From an evidence-design perspective, the useful practice is to keep factual measurement materials complete and contemporaneous, and to let counsel decide which communications sit in a privileged channel. Retroactive attempts to characterize an incomplete file as privileged rarely repair the underlying gap.

A minimum trail that holds up under questions

In practice, a defensible trail looks ordinary when it is done on time. There is a signed scope. There is a data extract with a dictionary and coverage statistics. There are reproducible metric outputs with cell counts. There is a clear statement of auditor independence where required. There is a version pin for the system audited. There is the published or reported output where publication is required. There is a findings register with management responses and dated closure evidence. Retention follows counsel's instruction and the organization's record policy, rather than an auditor's inbox retention defaults.

None of this substitutes for sound measurement. Poor metrics, carefully filed, remain poor metrics. The reverse is also true: sound metrics with no file force counsel to recreate the audit under deadline, often without the people who ran it. Boards asking whether an audit was "done" should ask a sharper question: if we had to explain this finding to a regulator or opposing expert in ninety days, which documents would we hand over, and would those documents tell one coherent story?

That is the standard the evidence trail is built to meet. The audit's conclusions matter. The audit's reconstructability is what lets those conclusions continue to matter after the engagement ends.