top of page
Search

Your Bank Has AI Agents - Who Governs the Agents, and Who on Your Board Has Sufficient Background to Question Them?

Your AI Agents Need The Discipline That Your Other Models Already Have

I have spent more than 35 years watching financial institutions get into serious trouble, rarely because their controls were missing, almost always because either nobody was watching them, or signal escalation failed.


The same pattern is now playing out with AI, in every industry, at a scale boards are only beginning to grasp.  And yet, a 2026 survey found that 54% of boards have not placed AI govern

ance in their top five priorities.


Right now, somewhere in the organization for which you are a Director, an AI agent is making or providing critical input into a risk decision. Something you may have approved in a board meeting – buried in a deck – six months ago and have not thought about since.


An AI agent may be analyzing a fraud alert, decisioning a credit application, analyzing documents, requesting missing data from a customer, conducting compliance or risk tests, updating operating procedures for regulatory changes,  and other consequential tasks. 


🚩Does anyone know exactly what your bank’s AI agents are doing, and do you as a director know who is supposed to be checking?

🚩Does your board need a data science expert in the same way Sarbanes-Oxley said that it needs a financial expert?


Financial services built an answer to similar questions after regulators forced the issue in 2011. As of February 2026, that answer has been rewritten specifically for AI, with Treasury's name on it. Most institutions have not noticed, and still fewer have implemented it.


In this post, I present why model validation discipline is the right frame for AI governance, what the US Treasury’s new framework gives you, and what boards should be asking for now.  


The Contractor and the Machine


To understand why model validation matters, go back to before 2011, when detective controls like transaction monitoring were a far more primitive exercise than they are today.


In those days, monitoring meant thresholds. Numeric cutoffs that triggered an alert when a transaction exceeded a dollar amount or frequency threshold. That much has not completely changed. What changed is the governance around those thresholds. Before 2011, there was not much.


I remember a client. A bank operating under a significant cease-and-desist order in the early 2000s. While we were on-site, a detail surfaced that captured everything wrong with the era. A contractor was quietly adjusting the thresholds in their transaction monitoring system and had been doing so for an extended period. Nobody could explain his methodology. Nobody had reviewed his work. Nobody was watching. Albeit well-meaning, he was just making changes, and the institution was none the wiser.

The board and senior management had no idea. They assumed, as is natural, that the existence of a control meant the control was working.


That assumption was wrong, and it cost them dearly.


This was not unusual. It was the industry standard.


The Birth of Model Validation


The regulatory response was OCC Bulletin 2011-12 and the Federal Reserve's SR 11-7. Together, they changed how institutions were required to think about any process driven by a model.


What made them remarkable was their reach. Regulators took validation principles built for credit risk and market risk models, the complex quantitative tools banks had run for decades, and applied them broadly to non-financial risk and compliance processes. The logic was simple. If a process uses a model to make critical decisions, that model must be understood, documented, challenged, and monitored. It does not matter whether it sits in risk management or in compliance.


So What Is a Model?


Under OCC 2011-12/SR 11-7, a model is any quantitative method, system, or approach that applies statistical, economic, financial, or mathematical theory to transform input data into estimates for decision-making.

That definition is deliberately wide. It captures the transaction monitoring system in your compliance department as surely as it captures a credit scoring engine.

With that net cast, institutions were on notice. Your compliance systems are models. Models require governance: inventory, documentation, independent validation, ongoing monitoring, formal challenge. That is the framework the industry spent the next decade building.

It is also the framework your AI agents are currently operating outside of.


Non-Deterministic Behavior: Where the 2011 Framework Runs Out


OCC 2011-12 and SR 11-7 were never meant to be static. It set a principle, not a technology standard, and the principle has held up. But AI strains it in ways the 2011 drafters did not anticipate.


Traditional models were static. You could test, validate, and monitor them periodically. AI models don’t always work that way.


Some AI models learn. They adapt. They behave differently next month than they did last month, based on data they encountered along the way. Periodic validation was never designed to catch that.


Some AI models, like the LLMs (e.g., ChatGPT, Claude, Gemini, DeepSeek, etc.) powering your AI agents, are “Non-deterministic,”  meaning they can generate different outputs from the same input.  According to recent academic research, “Experiments reveal accuracy variations of up to 15% across runs, with a gap of up to 70% between best possible performance and worst possible performance. No LLM consistently delivers the same outputs or accuracy across tasks.” 


Why this is a burning platform


The use of these non-deterministic tools is growing quickly and getting infinitely more complex.


According to a recent Deloitte survey, nearly 3 in 4 (74%) companies plan to deploy agentic AI within two years. IDC estimates this as a 10x increase by 2027. Additionally, Gartner reports that 40% of enterprise applications will be integrated with task-specific AI agents by the end of 2026, up from less than 5% one year ago.  Despite this growth, frighteningly, only 1-5 (21%) companies in the same Deloitte survey currently self-report a mature model for governance of autonomous agents, given the technology’s rapid adoption trajectory.


What’s more, in the growing world of AI agent deployments, we are not just dealing with the inconsistency of single LLMs and their newly trained versions.  Architectures such as LLM Mesh and internally constructed options deploy agents that use multiple swappable LLMs, thereby multiplying complexity.


Add to this an explosion in the unauthorized use of artificial intelligence tools, applications, or models by employees without the knowledge, approval, or oversight of their IT and security departments.  This so-called Shadow AI has been reported in 70% of organizations, with a 36% year-on-year increase.


Mixed Signals from Washington and Tougher Lines Abroad


US regulators- notably the Fed and the OCC- have been inconsistent on the near-term answer.  In its revised model risk management guidance, issued in April 2026, the footnote on page 3 is as follows:


“Generative AI and agentic AI models are novel and rapidly evolving. As such, they are not within the scope of this guidance. Nonetheless, a banking organization’s risk management and governance practices should guide the determination of appropriate governance and controls for any tools, processes, or systems not covered in this document. However, the principles described in this guidance apply to traditional statistical and quantitative models and non-generative, non-agentic AI models.” [emphasis added] Apparently, you are on your own to divine your own “appropriate governance and controls.”


Other jurisdictions are more explicit. The UK's SS1/23 was the first banking regulation written specifically for the characteristics of AI and machine learning models, and the EU AI Act reaches any AI system that could significantly affect health, safety, or fundamental rights, in any industry.


Other industry regulators in the US have the message too.  For example, the FDA's 2025 draft guidance on AI-enabled device software requires a Total Product Life Cycle approach: model documentation, data lineage, performance validation, bias analysis, continuous monitoring. Read that list again. It is the old SR 11-7 wearing a lab coat. 


The Framework May Be Catching Up


Here is what changed, and why I am writing this now:  there is a growing risk and an apparent vacuum of clear, actionable guidance.  As a result, I am concerned that Boards and their Risk and Audit committees underappreciate their current risks, and the control infrastructures that their financial institutions should be building.  


Help on where to start has already arrived but has received little attention in financial services risk management outside traditional cyber-risk circles. 

In February 2026, the Cyber Risk Institute released the Financial Services AI Risk Management Framework, the FS AI RMF, developed with the Financial Services Sector Coordinating Council and the U.S. Department of the Treasury. Treasury did not merely cite it. Treasury co-developed it and released it, alongside a shared AI Lexicon, as part of the federal AI Action Plan. 


“This work demonstrates that government and industry can come together to support secure AI adoption that increases the resilience of our financial system.” Secretary of the Treasury Scott Bessent. 


More than 100 institutions contributed: community banks, credit unions, national and multinational banks, insurers, with technical input from NIST.


This design matters. The FS AI RMF was built as a complement to risk management standards, including model risk management, not a replacement. It aligns structurally with the NIST AI Risk Management Framework and extends it through 230 control objectives covering governance, data, model development and validation, monitoring, third-party risk, and consumer protection. Those objectives are organized around four adoption stages, so a community bank running a chatbot is not held to the same intensity as a G-SIB running autonomous agents.


The discipline this post argues for is no longer something financial institutions have to invent and wait for the OCC and Fed examiners to cite. It now has a name, a structure, and Treasury's backing.


Activity Is Not Governance


Many institutions take comfort in the existence of AI policies, committees, steering groups, and board reporting.


But a policy is not a control, a committee is not a validation, and not all reporting is equally probative.


The evidence is specific to our industry, not anecdotal. Deloitte's 2026 assessment of 135 banks, including G-SIBs and D-SIBs, found 87% of global banks believe that they have clear room to improve their AI governance. The weakest dimension was organizational structure: who actually owns this. Publicly reported AI incidents in financial services rose nearly eightfold between 2022 and 2025. Our sector's share of all reported AI incidents more than doubled, from 1.8% in 2021 to 5.2% in 2025. And governance for agentic AI, the autonomous systems your board is approving right now, lags furthest behind.


International supervisors are converging on the same message. The Financial Stability Board assessed AI's financial stability implications in 2024, and in June 2026 issued sound practices for responsible AI adoption, flagging model risk, data quality, third-party concentration, and governance as priority vulnerabilities. The direction of travel is not ambiguous.


What Boards and CROs Should Do Now


That contractor at my client's bank, quietly adjusting thresholds with no documentation, no oversight, and no challenge, is not a historical footnote.

He is the template for what happens at scale when powerful models control agents that perform critical processes without governance.


Agentic governance gaps do not announce themselves. They accumulate quietly, and they surface as regulatory sanctions, litigation, and harm to real people, long after the window to prevent them has closed.


Here is a practical sequence. For those of you in financial services, the FS AI RMF provides an instrument for each step rather than a blank page.


  1. First 30 days: Ask management for the AI inventory, including all deployed AI Agents. Every AI system or agent that touches customers, drives decisions, manages workflows, or exercises real autonomy, classified by risk. The FS AI RMF's adoption stage questionnaire does exactly this across six dimensions. If management cannot produce an audit-validated picture, you have found your first governance finding.

  2. First 60 days: Establish AI governance following the guidance of the FS AI RMF.  This should include all 81 (!) specific governance controls outlined in their guidance.

  3. First 90 days: Establish that no AI deployment enters a critical process without independent validation. Data quality, model logic, performance testing, bias analysis, explainability, challenged by someone who did not build it. The framework's control objectives tell you what “validated” actually means for an AI system, and what evidence to expect.

  4. First 6 months: Put ongoing monitoring in place for every deployed system, with documented escalation thresholds and human review. A model that passed validation last year is not a model you can forget. AI drifts. The world moves around it.

  5. By 12 months: Assess whether your committee structure is honest. Audit committees are already carrying a heavy load. Ask whether AI model risk is genuinely being governed there, or merely reported there.


Note what these questions have in common. Not “did we validate this before launch?”

The question is whether someone is watching it today.


So does your board need a data scientist?


In the same way that the Sarbanes-Oxley Act of 2002 required company boards to name an “Audit committee financial expert,” companies and boards should consider whether they have members with sufficient expertise to govern Agentic AI-related risks. Once your board has a good feel for the nature and extent of its AI risks and controls, the answer should be evident.


Even without a bona fide expert, boards should assess whether they have directors who can ask precise questions about Agentic AI risks, recognize incomplete answers, and hold management accountable when governance is absent. That takes real knowledge of the current state of agentic AI, and engagement with what AI is doing in your enterprise. Not presentations about AI strategy. Substantive discussion of accountability, model risk, validation status, and monitoring results.


Do Not Wait for the Examiner


In 2011, the framework arrived as a mandate, after things had already gone wrong. This time we have the instrument before the mandate. The FS AI RMF is voluntary guidance today. It also carries Treasury's backing, and it is precisely the kind of framework supervisors convert into examination expectations. Institutions that align early will be ready when expectations harden. The rest will be explaining themselves.

The absence of a rule is not the absence of a risk.


Final Thought


Financial services did not build model validation because we were virtuous. We built it because regulators forced the issue after governance failures caused real harm.

We had our contractor-in-the-machine moment.


Your board now has the chance to get ahead of its own. The tools exist. The framework is written. The only question is whether your institution applies it before the cease-and-desist arrives, or after.


If you have questions, please get in touch with me at: jeff.lavine@jpladv.com.



 
 
 

Comments


  • LinkedIn
image.png

Subscribe to Our Newsletter

Contact Us

bottom of page