- Open Banker
- Posts
- Test AI, Against What?
Test AI, Against What?
Written by Delicia Hand
Delicia Reynolds Hand is the Senior Director of Digital Marketplace at Consumer Reports. With 20+ years at the intersection of technology, policy, and social impact, she leads initiatives evaluating fintech apps and AI's impact on financial services while developing frameworks for responsible implementation.
Open Banker curates and shares policy perspectives in the evolving landscape of financial services for free.
Consumer Reports buys the cars it tests. Not from the manufacturer, and not at a discount. Our shoppers walk onto the lot, negotiate like anyone else, and pay retail. The dealer usually finds out afterward who was buying. One of them told our test track staff what they all seem to think: had he known where the car was going, he would have treated the customer differently.
This is the reason we do what we do, how we do it. We spend more than thirty million dollars a year buying the things we test, sending shoppers to stores across the country, and refusing free samples and advertising. The method is ninety years old and it rests on one idea. A product handed to you by the company that made it is not the product a regular person actually gets.
Neither Impartial, Nor a Spectator
I have been thinking about that dealer while reading as much as I can of what people are writing about how to govern AI in financial services.
Consider who evaluates these systems today. The bank validates its own model, against standards it interprets. The vendor attests to its own governance and sells the attestation. The auditor is engaged by the firm and scoped by the firm. The examiner arrives announced. Every one of these is a serious exercise conducted by serious people, and every one is conducted by someone the system knows.
Software and intelligent systems raise the stakes in a way a refrigerator never did. A car cannot tell that it is on a test track. Large language models, however, can be quite good at recognizing when they are being evaluated, and researchers have documented systems that behave differently once they infer they are being watched. But the more basic problem requires no such sophistication. A product demonstrated in a conference room, on a curated prompt, to an audience that knows what it is looking for, is not the product that answers someone at eleven at night when their rent is due. Who sits in their chair?
Learned Experience
Sit in it and you find things. A woman applies for a car loan in her bank's app and gets an answer in seconds, and the answer is no, and the reason given is her credit profile. The message does not say which part of her credit profile. It does not say that the law entitles her to that specificity. It does not tell her that if the information behind the decision is wrong she can dispute it, or that she can ask for a person to look at the file again. She reads it twice, closes the app, and goes looking for a worse loan somewhere else.
Nothing about that requires bad faith. Her notice may well have been reviewed before it reached her. Many banks review adverse action language for sufficiency under Regulation B, monitor complaints, and run fair lending analysis on their underwriting. That work is real, and the people who do it are not asleep. It just was never designed to answer the question she was left holding, which is what she is supposed to do next.
Furthermore, the apparatus built to inspect these systems is pointed elsewhere. In April, the federal banking agencies replaced fifteen years of model risk guidance with SR 26-2, which defines model risk as the risk of financial loss, errors in financial statements and reporting, and flawed financial and risk management decisions. Every injury on that list is an injury to the institution, and scrutiny scales with how much of the bank's business rides on the model. This is a coherent way to write prudential guidance and I have no issue with it. It is also, by its own footnote, a document that places generative and agentic AI outside its scope, on the grounds that they are novel and rapidly evolving. The systems that talk to consumers, explain decisions to them, and increasingly move their money are the ones now sitting outside banking’s most developed governance regime. The one part of the bank a customer ever touches is the part nobody has been assigned to check on her behalf.
Marking the Bench
The gap here is not one of law, and not one of guidance. It is a gap in evidence, and you cannot collect evidence until you have settled what you are collecting it against. Our engineers cannot rate a car seat without a definition of what the seat must do to the crash forces in a thirty-five mile per hour collision. Point us at every financial AI on the market tomorrow and, without criteria, we would come back with anecdotes.
The people trying hardest to measure these systems have arrived at the same difficulty from another direction. In security, researchers reported recently that the benchmarks used to gauge what frontier models can do have been outrun by the models themselves. The old tests posed staged, isolated puzzles; the systems now act across long sequences in messy environments, and Stanford’s AI Index warns that evaluations built to last years are exhausted in months. Their question is what a model is capable of. Ours is what a product does to the person using it. Both questions fail the same way when they are asked once, in a clean room, of software that knows it is being asked.
The consumer criteria did not exist. So we wrote them.
Introducing the Consumer Finance AI Standard
The Consumer Finance AI Standard we published last month sets nine principles against a single question a consumer can hold any product to: whether it is actually working for him or her. It sets expectations against manipulation, against decisions no one can contest, and against the conflicts that surface when a product serves the company rather than the customer. Alongside each principle we published the criteria and the evaluation procedures, because a principle no one can check is not a standard. It does not substitute for model risk management, and it does not substitute for a certification body. It covers what those, by design, leave out, and it is written so that someone other than us can pick it up and run the test.
One of the nine principles is a duty of vigor, and it is the principle you can only see from the chair of the woman shopping for an auto loan. Nothing learned by inspecting a model from the inside tells you whether it mentioned to her that the denial was too vague, that the fee was disputable, that she had a right to human review. The duty of vigor asks a product not merely to avoid harming a consumer but to act for her when it sees her rights in play, in the moment, in plain language, without her having to know the statute or think to ask. No technological limit prevents this. A system that can compose an adverse action notice can flag everything else the consumer should know or think to do. Products do not do it because firms rarely gain when their customers push back, and no practice becomes standard through voluntary loss.
That arithmetic is going to change, which is why the criteria need to exist now rather than after the next round of guidance. Give a consumer an agent of her own and she “reads” the fine print every time, in milliseconds. She compares the offer against the market before she accepts it. She notices the fee, and she notices when the denial is vague. Disloyalty that survives today because it is tedious to detect becomes, in that market, a measurable property of a product, and measurable properties become competitive ones. What firms will need is something to be checked against, and consequences that make the checking count. That is the argument for a consumer standard, and against relying on voluntary commitment, which has never been entirely sufficient and will not start being sufficient because the software got better.
A financial product is not a refrigerator, and we do not test it by putting it on a loading dock. We download the app. We open the account, fund it, move small amounts of money through it, and read what it tells us when something goes wrong. We sit through the flow the way a customer sits through it, and nobody at the company knows which account is ours. That is the same method as the car, applied to a thing that has no doors.
We do not intend to test this market alone, and we should not be the only ones who can. That is why the evaluation procedures are published alongside the principles. A lab can run them. A bank can run them against its own product before it ships to see if it meets the consumer standard, which is the cheapest place anyone will ever learn if the answer is no. A regulator can cite them. We are working now with labs and with financial institutions to put them to use, and to sharpen them where the market shows us we have gotten something wrong.
She closed the app and went looking for a worse loan. The question the whole standard turns on is whether anyone was ever going to check the product from where she sat. We are. And we have written down how so others can do the same.
The opinions shared in this article are the author’s own and do not reflect the views of any organization they are affiliated with.
Open Banker curates and shares policy perspectives in the evolving landscape of financial services for free.
If an idea matters, you’ll find it here. If you find an idea here, it matters.
Interested in contributing to Open Banker? Send us an email at [email protected].
