AI Interviewer, audited for bias. Published in full.
Prepared by BABL AI Inc. | Audited June 29, 2026 | Signed: July 28, 2026
Audit Results · NYC Local Law 144
Audit conclusions
Three sections audited. Three passing opinions.
The audit was conducted by BABL AI Inc., an independent auditing firm whose lead auditors are ForHumanity Certified under the NYC AEDT Bias Audit standard. BABL AI independence conforms to the ForHumanity and Sarbanes-Oxley definitions. Fees are unrelated to the opinion rendered.
The system
What AI Interviewer does
AI Interviewer conducts AI-powered interviews. Given a job description, the AI Interviewer model generates a list of interview questions and evaluation criteria for the role for a customer’s review, revision, and/or approval. If the model determines that a candidate did not fully answer a question, it asks a follow-up so the candidate can give a complete answer.
After the interview completes, the model helps assess the candidate against the customer’s defined evaluation criteria, to help inform recruiters’ and hiring managers’ hiring decisions.
Methodology
How scoring rates and impact ratios work.
The audit used the scoring rate method – the proportion of candidates within a demographic group who scored at or above the overall median score of the full population. Impact ratios are calculated by dividing each group’s scoring rate by the scoring rate of the highest-scoring group. Under the federal Four-Fifths Rule (UGESP, 1978), an impact ratio below 0.80 is generally regarded as evidence of adverse impact.
In plain terms: male candidates scored at or above the median 0.754 of the time and female candidates 0.747 of the time, giving female candidates an impact ratio of 0.990, well above the 0.80 threshold. All groups in this audit remained above 0.80.
Data Transparency
Where these numbers come from.
Gender
Data from real historical use, with self-reported gender labels.
Age
Data from real historical use, with inferred age labels.
Race/ethnicity and intersectional
Synthetic data. Because insufficient real-world data was available for meaningful race/ethnicity and intersectional testing, a large language model was instructed to act as a candidate and conduct synthetic interviews across 12 candidate personas. Demographic data was incorporated through candidate names generated for each demographic group.
DISPARATE IMPACT · GENDER
Gender scoring rates, from real interviews.
| Group | N applicants | Scoring rate | Impact ratio |
|---|---|---|---|
| Male | 1,250 | 0.754 | 1.000 — reference |
| Female | 395 | 0.747 | 0.990 |
An additional 1,235 applicants with an unknown gender category were not included in this calculation, a reflection of real-world self-declaration gaps.
DISPARATE IMPACT · RACE / ETHNICITY
Race and ethnicity scoring rates across seven groups.
| Group | N applicants | Scoring rate | Impact ratio |
|---|---|---|---|
| Black or African American | 238 | 0.676 | 1.000 — reference |
| Two or more races | 238 | 0.672 | 0.994 |
| Native Hawaiian or Pacific Islander | 238 | 0.664 | 0.981 |
| Asian | 238 | 0.660 | 0.975 |
| Hispanic or Latino | 238 | 0.660 | 0.975 |
| White | 238 | 0.655 | 0.969 |
| Native American or Alaskan Native | 238 | 0.651 | 0.963 |
All seven race/ethnicity groups showed impact ratios above 0.80.
Disparate Impact · Age
Age scoring rates, from real interviews.
Real historical usage data. Age labels are inferred.
| Group | N applicants | Scoring rate | Impact ratio |
|---|---|---|---|
| Above 40 | 121 | 0.802 | 1.000 — reference |
| Below 40 | 2,759 | 0.723 | 0.902 |
Disparate Impact · Intersectional
Gender and race/ethnicity combined.
New York City Local Law 144 requires intersectional analysis of every combination of gender and race/ethnicity. The gender data here does not refer to the same candidates as the gender table above, and is synthetically constructed. The reference group is Black or African American Male.
Male candidates
| Group | N applicants | Scoring rate | Impact ratio |
|---|---|---|---|
| Black or African American | 119 | 0.689 | 1.000 — reference |
| Native Hawaiian or Pacific Islander | 119 | 0.681 | 0.988 |
| Two or more races | 119 | 0.672 | 0.976 |
| Asian | 119 | 0.664 | 0.963 |
| Hispanic or Latino | 119 | 0.664 | 0.963 |
| White | 119 | 0.655 | 0.951 |
| Native American or Alaskan Native | 119 | 0.647 | 0.939 |
Female candidates
| Group | N applicants | Scoring rate | Impact ratio |
|---|---|---|---|
| Black or African American | 119 | 0.664 | 0.963 |
| Native Hawaiian or Pacific Islander | 119 | 0.647 | 0.939 |
| Two or more races | 119 | 0.672 | 0.976 |
| Asian | 119 | 0.655 | 0.951 |
| Hispanic or Latino | 119 | 0.655 | 0.951 |
| White | 119 | 0.655 | 0.951 |
| Native American or Alaskan Native | 119 | 0.655 | 0.951 |
All intersectional groups showed impact ratios above 0.80. The lowest observed ratio was 0.939.
Governance · Audit finding: Pass
Who owns fairness at Eightfold AI.
Governance of bias and fairness risk is managed by a cross-functional Responsible AI working group that includes the Chief AI Compliance Officer along with product, engineering, legal, and security representatives. The audit confirmed that an accountable party is identified, that its duties are clearly defined, and that those duties were carried out prior to the audit.
Accountable party: Legal Department
Contact: legal@eightfold.ai
Risk Assessment · Audit finding: Pass
How Eightfold AI identifies and monitors bias risk.
The audit reviewed the internal risk assessment process, including the risk register, the risk prioritization methodology, and evidence of ongoing monitoring. BABL AI reviewed risk register dashboards, meeting minutes, and testimony from the maintainers of the risk register for AI Interviewer. The risk assessment covered risk identification, stakeholder impact, severity, likelihood, risk sources, and controls.
Independent Auditor
Audited by an independent firm.
BABL AI Inc. is an independent AI auditing firm based in Iowa City, Iowa. Its lead auditors are ForHumanity Certified Auditors under the NYC AEDT Bias Audit standard. The BABL AI audit framework, the Criterion Audit Framework, is modeled after financial auditing practice and was published in the Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency. BABL AI independence is codified by the Sarbanes-Oxley Act of 2002 and the ForHumanity Code of Ethics. Fees are unrelated to the opinion rendered.
Signed by BABL AI Inc., July 8, 2026.
Scope and limitations
What this audit covers, and what it does not.
This audit was conducted to satisfy the bias audit requirement of New York City Local Law No. 144 of 2021. The audit does not certify that the model is bias-free. No audit can make that claim. It is not intended to demonstrate compliance with any legislation other than the NYC AEDT law. The report does not make any determination as to whether the model is, in fact, an automated employment decision tool under the law.
Assessed in this audit
gender, race/ethnicity (seven groups), age, intersectional groups (all gender and race/ethnicity permutations), internal governance, and risk assessment.
NOT Assessed
protected classes beyond race/ethnicity, gender, and age, including immigration or citizenship status, disability status, marital and partnership status, national origin, pregnancy and lactation accommodations, religion or creed, sexual orientation, and veteran or active military service member status.
The race/ethnicity and gender intersectional results are based on synthetic test data, as described above.
This audit covers AI Interviewer. Our talent intelligence products are audited separately: see the Matching Model bias audit results. →
Our commitment
We do this every year.
New York City Local Law 144 requires an annual bias audit. Yes, we conduct these audits because the law requires it, but we also go beyond because responsible AI is an ongoing practice, not a one-time event.
Questions about our methodology?
Contact our legal team for questions about our audit methodology, results, or responsible AI practices.
Curious about Responsible AI at Eightfold?
See how fairness and transparency shape every product decision we make — from model design to ongoing governance.