This is the second part, automated. Four things that are normally four separate projects become byproducts of scoring one applicant.
Months later, by someone reconstructing decisions from memory and a notebook nobody can re-run.
In a notebook, before launch, and then never again — because repeating it is a week of somebody's time each quarter.
An examiner asks about an applicant from March. Nobody recorded which model version scored them, so nobody can answer.
Dashboards that alert on p-values, which at scale means alerting on everything, which means alerting on nothing.
It is a systems problem. So the fix is a system, not a longer checklist.
These controls drive the deployed model, not an illustration. Every score here is written to the append-only ledger, because a system that claimed otherwise on its own marketing page would be lying about the one thing it is for.
There is no control here for race, ethnicity, sex, or age. The request schema has no field to put them in — enforced by a test, not by discipline.
Every decision is a row chained to the one before it by SHA-256. Append-only is enforced by database triggers, not by convention. Retrieval walks the chain from genesis, loads the exact model version the decision was made with, and re-scores the stored application.
Measured on the out-of-time year, which is both the larger sample and the stronger claim. These are live figures from the running system.
Under disparate impact doctrine, a business-justified model is still challengeable if a comparably effective, less discriminatory alternative exists. So the system searches for one, and records the space it searched.
Structured against the SR 11-7 validation pillars, cross-referenced to EU AI Act Annex IV and ECOA. Every figure is read from an artifact that some other command produced. Retrain the model, regenerate, and the document describes the new model with no editing.
A decline also produces a real adverse action notice, built from the ledger row rather than a re-score, carrying the ECOA notice verbatim and the specific principal reasons. It explains why the FCRA disclosures do not apply here instead of padding the letter with boilerplate that does not.
See both in the demoThe target is what the lender did, not what the borrower did. It learns to replicate historical approval behaviour, including any bias in it. Every disparity reported is a disparity in approval behaviour, not evidence about applicants.
Public HMDA excludes the single most predictive variable in consumer underwriting. Performance here is not comparable to a production model, and the apparent importance of the variables that remain is inflated because they stand in for the missing one.
Everything here is measured on a fixed historical extract. That establishes properties of a model artifact. It does not establish that a deployed system is compliant.
Finding no better alternative inside 226 challengers is not proof that none exists. The space searched is recorded as an artifact so the claim can be checked and extended rather than taken on trust.
Score applicants, retrieve and replay any decision, run a portfolio through the fairness engine, compare a challenger model, and generate the validation pack as a real PDF.
The guest session expires by itself. Accounts exist so that actions in the audit trail have a name against them — an override records who made it, which is the point of recording it at all.