Evolving model risk management in the age of AI

| Article

Over the past 15 years, banks have built robust model risk management (MRM) capabilities in three waves. First, they focused on regulatory and prudential models. Second, they expanded their activities to broader model governance, encompassing managerial models embedded in core banking processes. Finally, they embraced a new wave shaped by machine learning and large language models (LLMs), which enabled a shift toward AI, generative AI, and agentic AI use cases, including solutions that extend beyond traditional model definitions.

Today, more than 80 percent of global organizations have deployed gen AI or agentic AI in at least one business function, and adoption is accelerating. But in the financial industry, less than 30 percent of European banks have integrated gen AI and agentic AI models in their model risk management frameworks, according to the latest McKinsey EMEA Model Risk Management Survey (see sidebar “Survey methodology”).

More broadly, MRM continues to evolve. Successive waves of development mean many institutions have transformed MRM into comprehensive model governance frameworks, our survey shows. Their efforts are supported by rising supervisory expectations, deeper integration of models into business decisions, and the industrialization of AI. In Europe and the United Kingdom, evolution is being accelerated through updated MRM regulations, including the European Central Bank Guide to Internal Models 2025, the AI Act, and the Supervisory Statement SS1/23, alongside a growing scrutiny of third-party technology dependencies.

As models increasingly influence high-stakes decisions, from capital allocation to product pricing and now to client-facing, investment, and workforce-related use cases, the question is no longer whether they work as intended, but whether organizations can trust them to operate at scale. These models also need to demonstrate trust to supervisors and stakeholders and preserve accountability across increasingly complex AI value chains. MRM is therefore extending its role from safeguarding the business to enabling it to move faster while preserving resilience, explainability, and confidence.

The survey reflects institutional recognition of the need for scalable MRM. It shows that, over the past year, annual validation volumes have increased by more than 10 percent. Over the next year, about 80 percent of banks expect an increasing number of models to require validation. This pressure is particularly pronounced with AI, where the recent rise of gen AI and agentic AI has democratized the development of use cases across organizations, leading to a rapid increase in both the number and diversity of models and model-like solutions.

Leading institutions are thus repositioning MRM for AI as a strategic enabler for speed, resilience, and confidence. But in Europe in particular, MRM frameworks need to continue evolving in line with the progression from standard AI to gen AI and agentic AI. If banks can make that happen, the door is open to transition from AI experimentation to enterprise-wide implementation.

Building the foundations: Three MRM throughput levers

Effective MRM requires strong foundational capabilities, even before building new AI frameworks, which are essential for adapting to today’s increasingly competitive context and evolving regulatory landscape. European banks have worked over the past few years to improve scalability and efficiency. According to our survey, they are focusing on three levers to manage growing model volumes and complexity while maintaining efficiency and speed.

Scaling MRM to meet increasing demand

MRM demand is rising sharply, driven by continued growth in model usage across banking operations, including but not limited to the expansion of AI use cases. This shift is accompanied by evolving European supervisory and regulatory expectations, which go beyond traditional prudential models and increasingly address the broader governance of model-enabled decisioning.

In the European Union and the United Kingdom, the scope of MRM is changing materially. For instance, in the United Kingdom, the Prudential Regulation Authority recently implemented the Supervisory Statement SS1/23,1 setting out five principles, including a formal supervisory definition of “model,” which led to a broader model scope and strengthened expectations for robust MRM.

The novelty is therefore not an extension of traditional regulatory models but the expansion of MRM’s scope to new types of models and model-like systems, along with more granular and use-case-driven expectations. In our survey, 40 percent of banks identify regulatory complexity and compliance as a top challenge for the next 12 months, reflecting the need to respond to an increasingly multilayered European regulatory environment.

In response, banks are evolving their MRM frameworks in targeted ways. This includes broadening model definitions and inventory scope, reinforcing tiering mechanisms (for instance, moving from periodic to trigger-based validations for lower tiers), strengthening approval and governance processes, and reviewing monitoring and validation frameworks to address the specific risk requirements of new model scopes.

Automation and industrialization of validation activities

Banks are accelerating automation across report generation, monitoring, and validation testing to streamline processes, improve efficiency, and support model life cycle agility. Model automation can reduce validation time by 30 to 50 percent and model development time by 20 to 40 percent. With AI and agentic capabilities, these impacts can be even higher (for example, automatically interpreting test results or conducting completeness checks against validation standards). More fundamentally, these capabilities enable a shift from automating tasks to reconfiguring parts of the model life cycle to increase end-to-end speed and consistency. Agentic automation, for example, can automate the end-to-end validation process into a pipeline of automated building blocks, from data extraction to report generation. (See sidebar “Model Excellence Workspace”.)

Banks are also enhancing their technology infrastructure and development environments to support data integration, effective code migration, and tool modernization.

Efficient operating model and streamlined life cycle

Leading banks are improving efficiency across the model life cycle by clarifying roles and streamlining the interaction mechanisms between model development and model validation, while maintaining separation of duties. Common efficiency drivers include service level agreements, standardized processes, and clearer accountability across development, validation, and monitoring. These changes enable more consistent, scalable, and coordinated processes as model inventories grow.

Operating models are also evolving through centers of excellence (COEs). Model development is increasingly centralized, with about 43 percent of banks surveyed reporting one or several COEs. Specifically for agentic and generative AI, European banks report a mixed operating model: 25 percent of survey respondents say that a central team owns the end-to-end development of AI use cases; 43 percent operate a hybrid model, where a central team prioritizes and supports gen AI use cases, while several teams across the bank do the development; and 32 percent run a decentralized model, where AI use cases are identified and developed by each area. This variety of operating models reinforces the need for clear governance standards, handoffs, and accountability across central and federated teams.

Five imperatives to future-proof MRM for AI

While foundations enable scale, AI requires targeted adaptations to the MRM framework. On the one hand, the number of AI use cases is rising fast, with widespread adoption in daily routines (for example, automated generation of emails or communications) and the reuse of the same underlying LLMs for many different model use cases. This leads to increasing complexity in identifying model uses and risks.

On the other hand, AI introduces new and increased risks, including explainability, bias, fairness, robustness, cybersecurity, and third-party risks, which must be addressed while enabling scale.

In Europe, 29 percent of banks have already updated their model governance to account for gen AI risks, the survey shows, while a further 36 percent are making progress, meaning 65 percent have either implemented or are implementing gen AI-specific model governance changes (exhibit).

Banks are taking actions to integrate gen AI in model risk management frameworks.

Below are five imperatives that will define success in evolving MRM to support AI adoption.

Imperative 1: Enhance model definition and inventory

One of the most structural MRM challenges with gen AI models sits in defining what a “model” is, and, therefore, what should sit within the scope of MRM and the model inventory. Many AI solutions are no longer single artifacts but ecosystems combining data pipelines, foundation models, prompt layers, rule-based components, and external APIs.

In this context, a purely methodological definition of models is no longer sufficient. Model definition must evolve toward a use-based perspective, through which inclusion in inventory is driven by the role of the solution in decision-making and its associated risk. This requires organizations to expand the scope of MRM beyond traditional models to include AI use cases, embedded tools, and third-party solutions that materially influence business outcomes.

A key implication is the need to distinguish among models, tools, and broader AI systems, particularly when relying on external providers. This includes defining which elements fall under MRM oversight and ensuring consistent identification, tracking, and accountability across the inventory. Any review should be aligned with recent supervisory guidance on the definition of model and scope of MRM (for example, Supervisory Statement SS1/23 in the United Kingdom).

Overall, a clear, consistent, and scalable model definition and inventory is a prerequisite for effective MRM in an AI-driven environment, enabling downstream processes such as tiering, validation, and monitoring to operate effectively.

European institutions are already moving in this direction: 32 percent of surveyed banks have updated their model definitions to account for gen AI risks, and a further 32 percent are making progress. This suggests that model definition is becoming one of the first areas of MRM adaptation for gen AI.

Imperative 2: Shift to use-case-based life cycle and approval

AI requires governance at the use-case level because the same model can support applications with different risk profiles.

Banks are introducing structured life cycle management—intake, risk assessment, approval, and monitoring—with clearer accountability across business, technology, and risk functions. Model life cycle management also empowers the first line for lower-risk use cases while maintaining strong oversight for higher-risk applications.

As use cases scale, ownership becomes harder—but more critical—to define, particularly for approvals, issue management, and monitoring.

This shift is also visible in development standards: As gen AI development becomes more decentralized, clear standards will be essential to maintain consistency; 14 percent of surveyed European banks already provide model development guidelines and standards to developers for gen AI use cases, while 43 percent are in the process of development.

Imperative 3: Adapt model tiering to AI

Traditional tiering frameworks typically classify models based on financial impact, regulatory criticality, and methodological complexity. While these dimensions remain relevant, they are no longer sufficient in the context of AI solutions. Tiering criteria must evolve to reflect additional drivers such as model autonomy, data sensitivity, and other risks—such as bias, explainability, conduct, or reputational impact—which can be as material as pure performance risk.

In practice, tiering becomes a central element of the MRM framework, directly informing approval processes and ongoing monitoring. It determines the depth of validation, the frequency and intensity of monitoring, and the appropriate escalation mechanisms. A single risk tier defined at the model level may therefore understate risk in high-impact deployments—particularly when the same model is reused across multiple use cases or evolves rapidly through retraining, configuration changes, or prompt updates.

Banks should move toward use-case-level tiering, adopting a risk-based and efficiency-oriented approach. In this context, tiering is not just a governance requirement but one of the most important design decisions for scaling AI safely while maintaining proportionate oversight.

Imperative 4: Evolve MRM and the challenge validation framework

The validation framework must evolve in both scope and approach to keep pace with the characteristics of AI models. Validation is becoming more use-case-driven, less static, and increasingly embedded across the model life cycle.

Model monitoring has become a principal component of this framework. AI models, unlike traditional models, can learn, adapt, or reference rapidly changing data. Thus, validation can no longer rely solely on point-in-time assessments. Instead, confidence must be continuously established through real-time post-deployment monitoring, complemented by point-in-time validations and escalation mechanisms to manage residual risk.

In addition, the characteristics of gen AI and agentic AI outputs (for example, unstructured, open ended, unpredictable) require new validation techniques. Leading institutions are expanding their tool kits, complementing classical statistical tests with techniques such as prompt sensitivity, explainability and repeatability diagnostics, fairness metrics, retrieval-augmented generation testing (RAG), and adversarial input testing. Validation is also increasingly integrated with application-level testing, such as behavioral and user acceptance testing (UAT), requiring tighter coordination between MRM and first-line technology teams.

As the number and complexity of AI use cases increase, speed becomes a critical factor not only for validation but also for how quickly use cases can be assessed, approved, and deployed. Banks are therefore evolving their MRM frameworks to enable faster, more streamlined decision-making across the end-to-end model life cycle, combining risk-based validation with efficient approval processes and continuous monitoring, while maintaining a proportionate level of scrutiny.

Gen AI–specific validation capabilities are still maturing. In Europe, 18 percent of surveyed banks have already developed specific validation standards and procedures to address new gen AI model risks, while 32 percent are making progress. Similarly, only 4 percent have built dedicated validation tools and infrastructure, although 43 percent are making progress. This highlights a material gap between governance ambition and industrialized validation capability.

From an organizational perspective, 25 percent of surveyed banks have already built dedicated validation teams for generative and agentic AI models, while 14 percent are making progress.

Imperative 5: Rethink AI risk management and governance

Effective AI model governance typically integrates MRM as one component of a “horizontal” risk framework and oversight model, which also includes cyber risk, IT risk, legal risk, and regulatory risk, particularly given the reliance on third-party models (opaque to banks), APIs, data providers, and plug-ins. Third-party dependencies introduce risks such as data leakage, biased outputs, security vulnerabilities, unclear accountability, and regulatory noncompliance. Therefore, managing AI at scale requires banks to move beyond MRM to address this wider set of risks through comprehensive AI risk frameworks and responsible AI frameworks.

To manage this complexity, leading institutions are adopting a coordinated governance model across the lines of defense, combining business ownership, centralized enablement, and independent oversight. In practice, the first line owns AI use cases and their associated risks, including design, deployment, monitoring, and incident management. A centralized capability nerve—typically, an AI platform or COE often structured as a 1.5-line—provides shared standards, tooling, and technical enablement, embedding guardrails and enabling consistent execution across use cases. One McKinsey study reveals that 90 percent of banks have adopted a centralized function or COE to manage AI.

The second line of defense, specifically on MRM, should take on a number of responsibilities, including the following:

  • ensuring compliance with AI and model risk expectations
  • defining model risk policies and standards
  • setting AI risk appetite and thresholds
  • determining which AI use cases classify as models and/or fall into the scope of MRM
  • defining the AI model risk taxonomy and risk tiering
  • setting approval requirements based on model materiality (including risk control requirements, acceptable use frameworks, and minimum test standards)
  • defining documentation and evidence standards
  • providing independent validation, challenge, and oversight (including extended dimensions such as data governance, explainability, transparency, bias/fairness, robustness, or human-in-the-loop oversight)
  • managing the AI model inventory (including the identification of critical models)
  • monitoring AI model risk across the portfolio and reporting AI risks to the board and senior management

Banks around the world are accelerating their AI strategies, moving from piloting to rollout across operations and business lines. In this context, MRM must play a vital role as the governance layer that enables first-line ownership, independent challenge, and safe scaling.

Institutions that successfully evolve their MRM frameworks for AI will not only strengthen their resilience but will unlock faster and more confident AI adoption in their organization and, in turn, a reliable source of value creation.

Explore a career with us