Blog · Bias & Fairness

Why AI screening is less biased than a human reading your resume

Decades of field experiments show resume screening is one of the most biased steps in hiring. Here is what the research says, and how a structured voice interview scored from a transcript removes the signals humans discriminate on.

Velma AI

10 min read

Key takeaways

Identical resumes get roughly 50% fewer callbacks when they carry a Black-sounding name instead of a white-sounding one — and that gap has not measurably narrowed since 1989.

Recruiters spend about seven seconds on a resume, which is far too little time for anything but pattern-matching on name, school, employer, and formatting.

Velma never scores a resume. Every candidate answers the same structured questions in a voice interview, and a separate AI scores the written transcript against role criteria defined before anyone applied.

Removing a signal is not the same as removing bias, so the analysis model is tested for proxy leakage and monitored for adverse impact under the four-fifths rule.

AI is a screening aid, not a decision-maker. Every hiring decision at every company using Velma is made by a person.

Most hiring conversations about AI start from the wrong premise: that human judgment is the neutral baseline and algorithms are the risky new variable. The evidence points the other way. Resume screening — the step AI is most often accused of corrupting — is already one of the most reliably biased processes in the labor market, and it has been for as long as researchers have measured it.

This article lays out that evidence, then explains exactly how our screening process differs and how we test whether it actually holds up.

The seven-second problem

Resume screening is not careful evaluation. It is rapid pattern-matching under time pressure. Eye-tracking research from The Ladders found recruiters spend an average of 7.4 seconds on an initial resume scan. In that window, a reviewer cannot assess judgment, communication, or problem-solving. They can only register surface features: the name at the top, the school, the logos of previous employers, the formatting, and whether the career path looks conventional.

Those surface features are precisely the ones that carry demographic information.

What the research actually shows

The strongest evidence comes from correspondence studies — experiments where researchers send employers pairs of resumes that are identical except for one signal, then count the callbacks. Because the resumes are matched, any difference in response is discrimination, not qualification.

Names. In the canonical study, Bertrand and Mullainathan sent nearly 5,000 fictitious resumes to job ads in Boston and Chicago. Resumes with white-sounding names received 50% more callbacks than identical resumes with Black-sounding names — a gap equivalent to eight additional years of experience. Improving the resume's quality helped white-named applicants significantly more than Black-named applicants, meaning better credentials did not close the gap. (Bertrand & Mullainathan, 2004)

It has not improved. Quillian and colleagues meta-analyzed every field experiment on hiring discrimination conducted in the United States between 1989 and 2015. They found no change in discrimination against Black applicants over 25 years, and only a modest decline for Latino applicants — through a period of enormous investment in diversity training and awareness. (Quillian et al., 2017, PNAS)

Gender. In a randomized double-blind study, science faculty rated an application for a lab manager position. The application was identical in every respect except the applicant's name. The male-named applicant was rated significantly more competent and hireable, and was offered a starting salary about $4,000 higher. Female faculty showed the same bias as male faculty. (Moss-Racusin et al., 2012, PNAS)

Age. A field experiment sending more than 40,000 applications found callback rates declining sharply with applicant age, with the steepest penalty falling on older women. (Neumark, Burn & Button, 2019, Journal of Political Economy)

"Fit" is often just similarity. Lauren Rivera's study of hiring at elite professional firms found evaluators explicitly weighting shared leisure activities, travel history, and self-presentation — treating cultural similarity to themselves as evidence of merit. (Rivera, 2012, American Sociological Review)

And when you remove the signals, outcomes change. When symphony orchestras moved auditions behind a screen so evaluators could hear but not see the musician, the probability that a woman advanced from a preliminary round rose substantially. Same musicians, same standard — different information available to the evaluator. (Goldin & Rouse, 2000, American Economic Review)

The orchestra result is the important one, because it isolates the mechanism. The problem was never that evaluators intended to discriminate. It was that they had access to information irrelevant to performance, and human judgment does not reliably ignore irrelevant information.

Why "just remove the names" doesn't work

The obvious fix — anonymize resumes — helps less than people expect. Demographic information leaks through proxies that are difficult to strip: the university, the neighborhood implied by an address, a graduation year that reveals age, the name of a sorority or a religious organization, gaps that correlate with caregiving, non-native phrasing, even the country a credential came from.

More fundamentally, an anonymized resume is still a document about a candidate's past access — which schools admitted them, which companies took a chance on them, whether they could afford an unpaid internship. It measures accumulated advantage at least as much as it measures capability. Redacting the name doesn't change what the artifact is.

If the goal is to evaluate what someone can actually do, the resume is the wrong instrument. It needs to be replaced, not cleaned.

What we do instead

Velma does not score resumes. There is no model that reads a CV and produces a ranking. The screening consists of two deliberately separated steps.

Step one — a structured voice interview. Every candidate for a role speaks with an AI interviewer that asks the same core questions, derived from the job requirements the hiring team defined before any applications arrived. Structured interviews with consistent questions and predetermined criteria are among the best-validated interventions in the personnel selection literature: they predict job performance substantially better than unstructured interviews and reduce the influence of evaluator bias. (Schmidt & Hunter, 1998, Psychological Bulletin)

Consistency also solves a problem no human panel solves: the fortieth candidate gets the same interview as the first. There is no fatigue, no drift in standards after lunch, and no candidate who benefits from a warmer conversation because the interviewer happened to share their hometown.

Step two — analysis of the transcript, not the person. A separate AI system evaluates the written transcript of the interview against the role criteria. That separation is the core design decision. The analysis model does not receive:

  • the candidate's resume, school, or previous employers
  • their name, photo, address, or contact details
  • the interview audio, so no accent, pitch, speech rate, or vocal quality
  • any demographic field, self-reported or inferred
  • how the candidate performed relative to anyone else

It receives the substance of what the candidate said and the criteria they are being measured against. What is left to score is the answer.

The screen in the orchestra audition worked because it removed information the evaluator did not need. Scoring a transcript against predefined criteria is the same idea applied to hiring: strip the channel that carries bias, keep the channel that carries evidence.

Being honest about how AI goes wrong

Any claim that AI is automatically fairer should be treated with suspicion, and the counterexample is well known. Amazon built an experimental resume-screening tool trained on ten years of its own hiring decisions. It learned to penalize resumes containing the word "women's" and to downgrade graduates of two all-women's colleges. The company scrapped it. (Reuters, 2018)

That failure has a specific cause worth naming: the system was trained to reproduce past human hiring decisions. It worked as designed. The bias was in the training target.

This is the failure mode we designed around. Our analysis model is not trained to replicate a company's historical hiring outcomes, and it does not learn from which candidates a company ultimately hired. It scores an answer against stated criteria. Removing the feedback loop from past decisions removes the mechanism that made Amazon's tool discriminatory.

The second real risk is proxy leakage: even without demographic fields, a model can pick up correlated signals from language itself — vocabulary, phrasing typical of a non-native speaker, regional idiom. Removing a variable is not the same as removing its influence. That risk cannot be designed away, only measured.

How we measure it

Fairness claims are only meaningful if they are testable. The measurement program:

Adverse impact monitoring. Selection rates are tracked across demographic groups using the four-fifths rule from the EEOC Uniform Guidelines on Employee Selection Procedures: if any group's pass rate falls below 80% of the highest group's, that is a flag requiring investigation, not an acceptable variance.

Counterfactual testing. The same transcript is re-scored with systematically varied surface features — names, gendered pronouns, phrasing patterns associated with non-native English. A fair scorer returns materially the same score. Divergence is a defect and is treated as one.

Criterion drift checks. Scores are audited against the stated role criteria to catch the model rewarding things nobody asked for, such as verbosity, confidence signaling, or vocabulary sophistication unrelated to the job.

Human-readable justification. Every score is accompanied by the reasoning and the transcript evidence behind it. A recruiter can see why, disagree, and override. An opaque score cannot be audited by the person relying on it; a justified one can.

People make the decisions

This matters legally and practically: Velma produces evidence, not verdicts. No candidate is advanced or rejected by the system. The hiring team reviews the transcript, the score, and the reasoning, and decides.

That design also aligns with where regulation is heading. New York City Local Law 144 requires annual independent bias audits and published results for automated employment decision tools. The Illinois AI Video Interview Act requires disclosure and consent. The Colorado AI Act imposes duties of care on developers and deployers of high-risk AI systems, employment included. The EU AI Act classifies employment screening as high-risk and mandates human oversight. GDPR Article 22 gives individuals the right not to be subject to solely automated decisions with significant effects.

Keeping a person in the decision seat is not a hedge. It is the requirement.

The honest comparison

The right question is not "is AI screening perfectly unbiased?" Nothing is. The question is whether it is less biased than what it replaces — a rushed human scan of a document that has been demonstrated, repeatedly and across decades, to produce large and stable disparities.

A structured interview scored from a transcript against criteria set in advance, audited for adverse impact and reviewed by a person, is a materially better instrument than seven seconds and a name at the top of a page. It is also a measurable one. You can audit an algorithm. You cannot audit an impression.

Frequently asked questions

Is AI screening less biased than a human reading resumes?

On the evidence, yes. Correspondence studies consistently find large, persistent bias in human resume screening — roughly 50% fewer callbacks for identical resumes with Black-sounding names, with no measurable improvement between 1989 and 2015. A structured process that asks every candidate the same questions and scores the written transcript against predefined criteria removes most of the signals that drive those disparities, and unlike human judgment it can be audited and measured.

Does Velma's AI make hiring decisions?

No. The AI conducts a structured interview and produces a score with written reasoning and transcript evidence. Every advance, reject, and offer decision is made by a person at the hiring company. This is a design requirement and it aligns with GDPR Article 22, the EU AI Act, and the Colorado AI Act.

What information does the scoring AI actually see?

The written transcript of the interview and the role criteria defined before applications opened. It does not receive the candidate's name, resume, school, previous employers, photo, contact details, the interview audio, or any demographic data.

Why analyze a transcript instead of the audio?

Audio carries accent, pitch, speech rate, and vocal quality — characteristics that correlate with national origin, gender, and age, and that have nothing to do with job performance. Scoring the text of what a candidate said removes that channel entirely.

Couldn't the AI still pick up bias from the language itself?

It could, which is why it is tested rather than assumed. The same transcript is re-scored with varied names, pronouns, and phrasing patterns associated with non-native English; a fair scorer should return materially the same result. Selection rates are also monitored against the EEOC four-fifths rule to catch disparities the counterfactual tests miss.

How is this different from Amazon's failed resume-screening AI?

Amazon's tool was trained to reproduce the company's own past hiring decisions, so it learned the bias in those decisions — including penalizing resumes that mentioned "women's." Velma's analysis model is not trained on a company's historical hiring outcomes and does not learn from who was ultimately hired. It scores an interview answer against stated criteria, which removes the feedback loop that caused that failure.

Are candidates told that AI is being used?

Yes. Candidates are shown a disclosure before the screening explaining that AI will analyze their responses, that a human makes the final decision, and how their data is handled. Details are in our privacy policy.

Sources

Last updated August 2, 2026

← All articles