Automation that cannot be evidenced is a liability, not an efficiency. What explainability actually requires in a credit decision — and the failure modes worth naming.
There is a version of AI adoption in lending that looks impressive in a demo and falls apart in a supervisory conversation. The model performs well. The metrics are strong. And then someone asks why a particular customer was declined, and the honest answer is that nobody can fully reconstruct it.
In most industries that is an inconvenience. In regulated lending it is a liability, and increasingly a reportable one.
The common framing treats explainability as a property of the model — something you bolt on with the right tooling once the thing works. That framing is why so many deployments stall at the governance gate.
Explainability in a lending context is not really about interpreting model weights. It is about being able to answer a much more practical set of questions, months after the fact, to someone who was not there:
None of those are model questions. They are process and record questions, and they have to be designed in from the start because the evidence has to be captured at the moment of the decision, not reconstructed afterwards.
Here is a position I have held since well before it was fashionable, and which I think has aged well.
An adaptive model that continuously learns is a genuinely uncomfortable thing to put inside a regulated credit decision. Given the same inputs, the outcome could differ depending on what the system has learned since. That may be entirely rational statistically. It is very difficult to defend to a customer, an auditor or a regulator — and it makes consistency testing something close to impossible.
The same application, assessed twice, should produce the same credit decision. A model that quietly changes its mind is a risk model you no longer control.
This does not mean rejecting machine learning. It means being deliberate about where non-determinism is acceptable. Ranking a collections queue by likelihood to pay is a reasonable place for it. Deciding whether a customer is approved is not — or at least, not without a deterministic policy layer sitting over the top and a human on the exceptions.
General caution is not a control. Agentic systems introduce specific risks that conventional model governance was not written for, and they are worth naming precisely because named threats can be designed against:
If your AI risk assessment does not mention any of these, it was probably written for a different technology.
Model accuracy tells you how the system performs when it works. It tells you very little about whether the deployment is safe.
More useful, and more likely to be what a supervisor asks about: how often does a human intervene, and is that rate moving? How quickly is a bad outcome detected? Does the escalation path actually function, or does it route to someone who no longer holds that role? What proportion of decisions could be fully reconstructed today if requested?
These are operational measures rather than data science measures, which is precisely why they tend to be missing.
None of this requires an AI risk committee, a model risk function or a multi-year governance programme. Most specialised lenders have none of those, and waiting until they do is not a strategy.
What a first deployment does need is proportionate: one clearly scoped use case, verified data sources, a defined point where automation hands over, a named person accountable for exceptions, an evidence trail produced as a by-product of the work, and agreement in advance on what would cause you to switch it off.
That is a manageable list. It is also, in practice, the difference between a deployment that survives its first review and one that quietly gets shelved.
Practical thinking on specialised finance, platforms and governed AI. Twice a month.