A 280-person insurance brokerage in Jersey City rolled out an AI resume screener in January. By March it was clearing 1,400 applications a week and shortlisting candidates in under an hour instead of the five days it used to take a two-person talent team. Everyone was thrilled. Then a candidate who got auto-rejected filed a complaint with the city, and the company discovered it had never run the bias audit that New York City's Local Law 144 requires. The tool worked beautifully. The compliance did not exist.
That gap — fast tooling, missing paperwork — is where most recruiting teams are living right now. The technology to screen at speed is cheap and everywhere. The discipline to prove it screens fairly is not. And regulators, starting with NYC, have decided that if you can't prove it, you can't use it.
What Local Law 144 Actually Requires
NYC Local Law 144 has been enforced since July 2023, and it applies to any employer using an "automated employment decision tool" (an AEDT) to screen candidates or employees for a position located in New York City. The trigger isn't where your company is headquartered. It's where the job sits. A brokerage in Jersey City hiring for a Manhattan office is squarely in scope.
The law has three concrete demands:
- An annual bias audit by an independent auditor, measuring the tool's selection rates across sex, race/ethnicity, and the intersections of the two.
- Public disclosure of that audit's results on your website before you use the tool.
- Candidate notice at least 10 business days before the tool is used, explaining what's being assessed and offering an alternative process on request.
Miss any of these and you're looking at fines of $500 for a first violation and up to $1,500 per violation per day after that. For a company screening 1,400 applications a week, "per violation per day" compounds fast.
The metric at the center of the audit is the impact ratio — the selection rate for each group divided by the selection rate of the most-selected group. The historical benchmark regulators lean on is the four-fifths rule: if any group is selected at less than 80% of the rate of the top group, that's a red flag worth explaining.
Why AI Screeners Fail Audits (When They Fail)
The failure is almost never someone writing "reject women" into a model. It's subtler and more mechanical than that.
Proxy variables. A model trained to predict "candidates like the ones we hired before" will happily learn that graduates of certain schools, ZIP codes, or gap-free work histories correlate with your past hires. None of those fields say "gender" or "race," but together they can reconstruct it. Amazon famously scrapped an internal recruiting model in 2018 after it learned to downgrade resumes containing the word "women's," as in "women's chess club captain."
Skewed training data. If your last 200 hires for a role were 85% men, a model that optimizes for resemblance to past hires inherits that skew and calls it accuracy.
Threshold effects. Two groups can have nearly identical average scores, but if you cut the shortlist at a hard rank — say, top 40 candidates — a small average difference can produce a large gap in who actually clears the bar. The audit measures who gets selected, not who scores well.
The uncomfortable truth is that a screener can be highly accurate at predicting your historical hiring and still fail a bias audit, because it's faithfully reproducing a biased history.
Human-in-the-Loop Is the Fix, Not a Formality
The phrase "human-in-the-loop" gets thrown around like a compliance sticker. Done properly, it's an actual control point that changes outcomes. Here's what it looks like in a recruiting pipeline that survives an audit:
- The AI ranks and explains; it does not reject. The tool produces a shortlist and, critically, a plain-language reason for each ranking ("8 of 10 required skills matched, 4 years relevant experience"). No candidate is removed from consideration by the machine alone.
- A recruiter reviews every auto-declined candidate above a threshold. If the model scores a candidate as borderline, a person looks before that door closes. This is both a fairness safeguard and, under 144's alternative-process provision, a practical necessity.
- Reasons are logged, not just scores. When the auditor arrives, "the model said 62" is useless. "Matched 6 of 9 must-have skills, missing the required license" is defensible.
- Override rates are tracked. If recruiters override the AI on 40% of a particular group's rankings, that's an early warning the model is miscalibrated for that group — long before the annual audit catches it.
The point of the human is not to rubber-stamp the machine's speed. It's to catch the cases the machine gets confidently wrong, and to create a record that proves a person was accountable for the decision.
A Compliant Screening Pipeline, Step by Step
For teams building or buying this, here's the sequence that keeps speed without the legal exposure:
- Inventory your tools. Anything that scores, ranks, or filters candidates counts, including features buried inside your ATS. Many teams fail 144 simply because they didn't know a vendor feature qualified as an AEDT.
- Run the bias audit before go-live, then annually. Use an independent auditor and test on data that reflects your actual applicant pool, not a vendor's generic benchmark.
- Publish the summary. Impact ratios by group, the audit date, and the distribution of the tool. On a public page. This is not optional.
- Send candidate notice. Ten business days out, describe what's assessed and how to request an alternative.
- Keep the human decision point. Every consequential rejection gets human review, and every review gets logged.
- Monitor between audits. Selection rates and override rates should be dashboarded monthly. An annual audit that surfaces a problem you've been running for eleven months is a bad day.
This is the model our AI for recruiting practice builds around: the screening runs at machine speed, but the rejection logic, the audit trail, and the human checkpoints are wired in from day one rather than bolted on after a complaint.
Speed and Compliance Are Not a Trade-Off
The Jersey City brokerage assumed it had to choose: keep the fast tool and accept the risk, or rip it out and go back to five-day manual screening. Neither was necessary. It ran an audit (impact ratios came back at 0.71 for one group, below the 0.80 line), retrained the model to drop three proxy features, added recruiter review on borderline declines, and published the corrected numbers. Screening time stayed under an hour. The complaint was resolved without a fine.
The lesson generalizes well beyond NYC. Illinois, Maryland, Colorado, and the EU AI Act are all moving toward the same posture: automated hiring decisions must be explainable, auditable, and ultimately supervised by a human. Building for NYC's standard today means you're mostly ready for the rest. For teams whose ATS won't support this natively, a custom AI software layer that logs reasons, tracks overrides, and generates audit-ready reports is usually cheaper than one round of fines and remediation.
The companies that win here treat the bias audit not as a tax on AI but as the thing that lets them use AI aggressively — because they can prove, to a regulator or a rejected candidate, exactly why each decision was made.
Frequently Asked Questions
Does Local Law 144 apply if my company is based outside New York City?
Yes, if the job itself is located in NYC or the candidate would work in NYC. The law follows the position, not the employer's headquarters. Remote roles based out of an NYC office are generally in scope, so check where the role is anchored before assuming you're exempt.
How often do we need a bias audit?
At minimum once every 12 months, by an independent auditor, and before you first use the tool. If you materially change the model or the data it's trained on, you should re-audit rather than wait for the annual cycle.
Can AI legally reject a candidate on its own?
It's not outright banned, but it's the fastest way to fail an audit and invite a complaint. Best practice, and the direction regulation is heading, is that AI ranks and recommends while a human owns the final rejection and that decision is logged with a reason.
What's the four-fifths rule?
It's a benchmark from EEOC guidance: if any demographic group is selected at less than 80% of the rate of the highest-selected group, the difference is presumed significant enough to warrant scrutiny. Bias audits report this as an impact ratio, where anything under 0.80 is a flag to investigate.
We already use an ATS with built-in scoring. Are we covered?
Only if that scoring feature has itself been audited and disclosed. Many teams don't realize a vendor's ranking feature qualifies as an automated employment decision tool. Inventory every scoring or filtering feature and confirm each one is covered, not just the standalone "AI screener" you bought on purpose.