SR 26-2 Model Risk Management

Meeting SR 26-2’s Demonstrable Evidence Standard for Model Risk Management

Share
Repost

SR 26-2 has drawn a mostly positive reaction, with some practitioners viewing it as easier to implement than SR 11-7. That’s a fair read given the real flexibilities built into the updated guidance:

  • A narrower definition of what counts as a model
  • Validation cycles set by materiality rather than the calendar
  • Independence is judged by the rigor of the review rather than reporting lines

But one requirement proves difficult for almost every institution trying to operationalize SR 26-2: the renewed emphasis on demonstrating that model governance, validation and oversight are supported by appropriate evidence and documentation. 

When a validation team seeks the clearest way to demonstrate rigor, benchmarking against an independent external reference can provide a particularly strong source of independent evidence. For portfolios with no rating and no market price, external benchmarking options are often limited. So, banks can describe their process, but can’t always produce the proof.

This guide walks through the three things SR 26-2 expects banks to demonstrate: effective challenge, continuous monitoring, and risk-based oversight, where the usual sources of evidence run out, and what closes the gap. First, a precise read on what the guidance changed, since that’s where most coverage blurs the line between flexibility and relief.

 

What SR 26-2 Changes

On April 17, 2026, the Federal Reserve, together with the OCC and FDIC, issued SR 26-2. This supersedes SR 11-7 as the primary reference for how large US banks manage model risk and incorporates SR 21-8.

The new guidance applies most directly to banking organizations with more than $30 billion in total assets. Smaller banks are generally excluded, unless they carry significant model risk, either because their models are unusually complex or because parts of their business fall outside traditional community banking. 

Four structural changes explain why the letter has been read as a loosening of the rules:

  • The definition of a “model” is narrower. To count as a model now, a tool must have a complex quantitative method, a basis in statistical, economic, or financial theory, and a quantitative output. Simple spreadsheet arithmetic and fixed rule-based systems no longer qualify. This shrinks the number of credit risk management solutions risk teams are responsible for tracking.
  • Validation no longer runs on a fixed schedule. SR 26-2 ties how often, how deeply, and in what way a model is reviewed to its materiality, how frequently it changes, and how much data exists to test it. In practice, this means a stable, low-stakes model no longer requires the same yearly review as a model that feeds into capital calculations.
  • Independence is judged differently. Independence is judged by how rigorous and effective the review itself was, not by the org chart. A bank can structure its teams more flexibly, as long as it can demonstrate that the challenge to the model held up.
  • Generative and agentic AI are excluded entirely. The guidance does not attempt to provide a comprehensive framework for rapidly evolving generative or agentic AI use cases, leaving firms to govern those under broader internal risk management frameworks.
  • SR 26-2 is limited to guidance, as it only sets expectations for sound practice. According to the release, deviating from it won’t, by itself, trigger supervisory criticism. Criticism arises when a gap in model risk management coincides with evidence of an unsafe or unsound practice elsewhere.
sr26-2 compliance

However, effective challenge for models still remains. The three pillars of validation (conceptual soundness, outcomes analysis, and ongoing monitoring) are unchanged. And banks are still fully accountable for any models they buy or license from vendors. What’s changed is what a bank must now prove. That breaks down into three supervisory expectations, each with its own built-in evidence requirement.

 

1. Effective Challenge Backed by Independent Evidence

Effective challenge survives SR 26-2 fully intact, and it’s still described as the central pillar of sound model governance. Someone with real expertise, independence, and authority has to scrutinize a model’s assumptions, design, outputs, and be in a position to actually force a change if something’s wrong.

Previously, a bank could point to only its process (annual validation, a review committee, a sign-off) as evidence. SR 26-2 goes beyond the process and asks for documentation showing what was compared, what the comparison revealed, and what the bank did in response.

In practice, a defensible answer proves that: 

  • The internal rating on a given counterparty is externally validated
  • Any gap between the internal and external rating is investigated rather than ignored
  • The whole trail (comparison, finding, and resolution) is documented well enough that an internal auditor or an examiner can follow it without a guide.

The most common failure is a challenge that’s procedural, not substantive, because the process was never actually capable of catching a problem, due to circularity. This means that whatever a bank uses to challenge a model must be genuinely independent of that model.

 

2. Continuous Monitoring That Detects Drift

Ongoing monitoring is one of the retained pillars. Its job is to ensure a model keeps pace with real-world changes, such as shifts in the portfolio, economic events, and changes in borrower behaviour. When any of these causes meaningful deterioration in the model’s performance, the guidance expects that finding to feed directly into a decision about recalibrating or rebuilding the model.

For the monitoring to be effective and acceptable, the guidance prescribes that:

  • Monitoring should be recurring, not a one-time event, and run between validation cycles
  • There should be a documented comparison of model output against something external to it
  • There should be defined thresholds for what counts as meaningful divergence and a clear escalation path when a threshold is breached.

This is the hardest pillar to evidence, for three reasons. First, annual backtests can’t catch drift that happens between cycles. Second, realized defaults are too rare in most wholesale portfolios to give a timely signal. 

By the time enough defaults accumulate to prove a model wrong, the damage is done. Lastly, the standard workaround, population-stability metrics, only detects when the inputs have drifted from the development sample, not whether outputs still rank risk correctly.

Additionally, whatever a bank monitors against has to refresh frequently enough to catch deterioration between validations. An external reference that updates only once a year can’t be applied to a portfolio that moves monthly.

 

3. Risk-Based Oversight Aimed Where Risk Concentrates

SR 26-2’s core principle is that the rigor of model risk management (validation depth, monitoring frequency, and governance sign-off requirements) should be commensurate with materiality. The more a model matters in terms of exposure and purpose, the more scrutiny it gets, ensuring resources are allocated where the risk lies.

For low-stakes models, it means that a pricing tool with minimal exposure might be validated every few years rather than annually, with a lighter review each time. But apply the same logic to your highest-stakes models, and the effect reverses. 

PD models feeding CECL provisions, regulatory capital, and credit decisions sit at or near the top of any risk ranking, which means they stay in the intensive-oversight bucket, no matter how much flexibility the rest of the framework allows elsewhere. In other words, the relief goes where the risk isn’t.

There’s a second obligation buried in “risk-based,” and it’s easy to miss: a bank that tiers its oversight has to be able to justify the tiering. Why did this model get more scrutiny than that one? Which sectors or counterparties warranted the extra attention, and how do you know? Answering these questions requires a signal showing where risk is actually building:

  • The sectors deteriorating
  • The counterparties other lenders are growing cautious on
  • Divergence between internal PD and an external benchmark

Whatever supplies that signal has to cover the exposures where the risk actually sits. For most wholesale books, that’s the private and unrated obligors that make up the majority of the portfolio. A reference that only covers rated names can’t tell you where risk is building in the part of the book that isn’t rated.

 

Why the Usual Evidence Falls Short

“Effective challenge” needs evidence that’s independent, “monitoring” needs evidence that’s frequent, and “oversight” needs evidence with real coverage. Run the sources most banks already have against those three tests, and each one fails at least one:

Agency ratings, even though they can pass for an independent benchmark, stop at the largest rated issuers, typically 10-15% of a wholesale book. Ratings also apply to entities at the parent level, while the exposure often sits with a subsidiary. Update cadence is typically every quarter or longer, which is too slow to catch drift between validation cycles.

Internal challenger models offer full coverage and can be built to refresh as often as needed. But a second model built by the same institution, often on the same data and the same assumptions as the first, isn’t an independent source.

Market-implied measures, such as CDS spreads, equity-based structural models, and vendor scores built on traded prices, are fast, forward-looking, and external benchmarks. However, they only exist where a market price or public financials exist, which is roughly the same rated, traded universe that agency ratings already cover. They’re also noisy, moving with sentiment as much as with credit quality.

When put side by side, each alternative has meaningful strengths, but also important limitations. None covers unrated entities, which means there’s a large blind spot that no combination of the usual tools fully resolves.

Closing The Coverage Gap With Consensus Data

Clearing all three tests requires an external benchmark that’s not tainted by bias or the opinion of a single entity. Consensus credit data, which is the pooled view of the many institutions that already carry the exposure, meets that criterion.

Consensus credit data works by collecting the same PD estimates that each contributing institution generates for its own book and aggregating them into a single reference figure per entity. To prevent bias, contributions are anonymized, and a minimum number of contributors is required before any figure is published. This is a genuinely different category from the three sources already covered: 

  • It isn’t a model, so it doesn’t inherit the circularity problem of a challenger built on the same assumptions as the model it’s meant to check. 
  • It isn’t a rating agency’s finished opinion either, which focuses only on entities large enough to justify the cost, so it doesn’t inherit their coverage limits. 
  • It’s built from many banks’ internal models, and pooling independent bank assessments creates an external consensus benchmark.

The difference becomes more apparent when you run it against the requirements established in the earlier sections. First, consensus comes from other banks’ own assessments, not from the model being checked, so it’s a genuine second opinion. It also refreshes weekly as banks submit new views, making it possible to catch a deterioration well before the next scheduled validation would. And because contributing banks lend to the same borrowers any given institution does, consensus reaches the private and unrated names that agency ratings and market-implied measures can’t.

 

Credit Benchmark as a Consensus Data Provider

Credit Benchmark provides an aggregated, regulator-supervised PD view across 40+ contributing banks, covering 100,000+ entities, around 90% of which carry no external rating, having over 10 years of history covering multiple credit cycles, regions, and industries. Mapped against the three requirements and the three areas of evidence built into SR 26-2, its value as an independent external benchmark becomes clear, with data being refreshed weekly.

First, “effective challenge” needs a genuine second opinion. Credit Benchmark supplies that because the consensus figure for any entity is built entirely from other institutions’ internal views, not from the bank’s own model or its inputs. State Street uses this capability to benchmark its internal ratings against Credit Benchmark’s peer consensus and uses Credit Benchmark to benchmark internal ratings against peer consensus as part of its credit governance process. to its enterprise risk committee, giving the challenge an external anchor that examiners can follow.

The next is continuous monitoring, which requires a signal that refreshes often enough to catch drift between validation cycles. Because Credit Benchmark is rebuilt on a rolling basis as contributing banks submit updated views, movement in the consensus figure shows up well before an annual or multi-year revalidation would. That’s why a global bank with over $1 trillion in assets uses it alongside agency ratings to flag divergence in funds and specialized finance between scheduled reviews.

Lastly, risk-based oversight needs coverage that reaches the exposures where risk actually concentrates, including names with no rating or default history to calibrate against. IRB Nexus, built with Oliver Wyman specifically for this problem, supplies bespoke default datasets and rank-order testing (Gini, Spearman, grade-level homogeneity), so a bank can direct oversight toward exactly the portfolios where an external check would otherwise be impossible. It’s available to contributing banks.

It’s worth pointing out that Credit Benchmark data is an input to a bank’s own effective challenge and outcomes analysis. It isn’t a model, and it isn’t a vendor model subject to validation under the guidance in its own right. Nor is it a replacement for internal models or agency ratings; rather, it’s simply an external benchmark that reaches the entities others don’t.

 

Getting Set For Your Next Validation Cycle

Effective challenge, continuous monitoring, and risk-based oversight were the standards before SR 26-2 and remain so. What changed is the burden of proof. A bank now has to document its process, not just describe it. And for the unrated majority of most portfolios, documentation has often been more difficult where external reference points are unavailable, particularly for private and unrated obligors. This is exactly the gap consensus data was built to close.

A practical starting point is knowing where the gap is. Credit Benchmark offers a coverage assessment that maps a portfolio against the consensus database, showing exactly which borrowers already have a Credit Consensus Rating.

Request a coverage assessment for your wholesale credit portfolio today.

Schedule a demo

Please complete the form below to arrange a demo.

    By submitting this form you agree to Credit Benchmark’s
    Privacy Policy and Terms and Conditions.

    Subscribe to our newsletter

      By submitting this form you agree to Credit Benchmark’s
      Privacy Policy and Terms and Conditions.