Skip to Content

Evidencing PD Estimates for Low-Default Portfolios Under the Internal Ratings Based Approach

Evidencing PD Estimates for Low-Default Portfolios Under the Internal Ratings Based Approach

Banks using the Internal Ratings-Based Approach are expected to produce robust probability-of-default estimates, but this becomes much harder when the portfolio itself has very few observed defaults. For exposures such as funds, hedge funds, and other non-bank financial institutions, internal historical data may be too sparse to give model validation teams a strong reference point.

The challenge, then, is not just how to estimate PD. It is how to evidence that estimate well enough to make it defensible. That is where external benchmarking becomes important. 

But not all external benchmarks provide the same kind of evidence. Agency ratings, vendor models, and peer-bank estimates differ in how they are produced, which populations they cover, and how closely they resemble the internal estimates being assessed.

For validation teams working with low-default portfolios, the more useful question is therefore not simply whether external data is available, but whether that data provides a relevant and credible reference point for the estimate under review.

This article examines how consensus PD estimates derived from banks’ internal models can help fill that gap.

Why low-default portfolios are difficult to validate

A credit backtest compares the defaults a model predicts with the defaults that actually occur. On a low-default portfolio, there may simply be too few defaults for that comparison to say much.

Take a portfolio of 150 fund counterparties with an average annual PD of 0.20%. Over five years, the model implies only about 1.5 defaults in total. If no defaults occur, that does not by itself show that the 0.20% estimate is wrong. Nor does it provide strong evidence that it is right. With so few expected events, the backtest has limited ability to distinguish between different low PD estimates.

why zero defaults doesn't mean the pd is right

The Basel Committee recognizes this limitation. It states that validation becomes less statistically reliable with few historical observations, which is one reason simple numerical pass/fail thresholds are difficult to apply to low-default portfolios. The ECB makes the related point that banks need to show that the data used for calibration is representative, and to justify any adjustments when it is not. 

The problem is especially relevant for funds, hedge funds and other non-bank financial counterparties. These exposures can generate little internal default history even in sizable portfolios. When that happens, validation teams explore three reasonable options:

Finding defaults somewhere else

One response is to enlarge the sample until there are enough defaults to test. The fund portfolio might be combined with a broader corporate population, or the observation window might be extended further back in time. Although this increases the number of observed defaults, those additional observations may not be representative of the portfolio being validated.

The problem here is representativeness. The Basel framework itself links the reliability of pooled PD estimates to the underlying borrowers being sufficiently homogeneous, while the ECB has stressed the need to assess whether calibration data are representative before relying on them.

Applying conservatism for estimation uncertainty

A more disciplined response is to add a conservative margin where the estimate can’t be evidenced. Basel expects conservatism precisely where estimation uncertainty is high, and it clears the immediate question from independent review. 

However, it does not resolve the problem for two reasons. First, if the margin itself cannot be supported by evidence, it simply adds another uncertain number to the PD. Second, a more conservative PD can increase risk-weighted assets. With the output floor already limiting the capital benefit of IRB models, additional conservatism can consume more capital without solving the original evidence gap.

Benchmarking against agency ratings

Agency ratings provide another established source of external challenge, particularly for public companies and large bond issuers. However, Agency ratings have three important limitations for this particular comparison.

First, an agency rating is expressed on the agency’s rating scale rather than as the bank’s calibrated one-year PD. To compare the two, the validation team normally needs a mapping from the agency scale to a PD scale. That mapping becomes another assumption that has to be justified.

Coverage creates a second limitation. Significant parts of the commercial credit universe (particularly private companies and some non-bank counterparties) may not have an agency rating at all. 

Provenance is another consideration. Agency ratings are typically issuer-paid, creating a potential conflict that validation teams must account for when assessing independence. The Bank of England’s Financial Policy Committee also identifies reliance on credit rating agencies as a vulnerability in private markets, alongside leverage, valuation challenges and links with banks and insurers.

Using a vendor PD model as the comparator

A fourth option is a commercial PD model based on financial statements or market data. These can work well as challenger models, but comparability is the issue: a point-in-time or market-driven PD is not directly equivalent to an IRB PD calibrated to long-run average default experience.

Coverage can also be limited where funds and non-bank financials do not publish the inputs those models require.

The common problem across these approaches is that a validation team can have more data, more conservatism, and more external information, yet still lack a genuinely comparable benchmark for the PD being tested. 

What each credit data source can and can’t do
Criterion Agency ratings Vendor PD models Peer-bank IRB estimates
Relevant coverage
Does it cover the funds and non-bank names in the book?
Many private companies and non-bank counterparties are unrated
Limited where entities don’t publish the financial-statement inputs required
Extends into private and unrated parts of the credit universe
Comparability
Same methodology, horizon and definition of default?
Ordinal rating grade, not a calibrated one-year PD — needs a mapping
Point-in-time or market-driven, not through-the-cycle — needs adjustment
PDs from other banks’ IRB models, under comparable supervisory expectations
Multiple independent observations
How many separate views sit behind the number?
One view per borrower
One view per borrower
Aggregated from multiple contributing institutions per name
Clear provenance
Can the validator explain where it came from?
Methodology is public, but issuer-paid — a vulnerability the FPC has flagged
Documented vendor methodology
Known contributor count and dispersion behind each estimate

Scroll the table sideways on smaller screens.

That is the gap any external PD benchmark must solve, to be considered useful for low-default portfolios.

What makes an external PD benchmark useful

When defaults are too scarce for backtesting to provide strong evidence on its own, external benchmarking provides a second line of evidence. Basel recognizes that banks may use external and pooled data alongside their own default experience when assessing PD estimates. The point is not to replace backtesting, but to strengthen the validation where internal evidence is thin.

But for that comparison to be useful, the external benchmark needs to pass four tests.

  1. Relevant coverage. A benchmark that covers only a small part of a fund or non-bank portfolio leaves the original validation gap largely intact.
  2. Comparability. The closer the external estimate is to the PD being tested (in methodology, horizon, and definition of default), the less adjustment is needed before the two can be compared.
  3. Enough independent observations backing the benchmark. A single external PD gives the validator one alternative view, while multiple independent estimates help the bank see whether its own estimate is broadly aligned or an outlier.
  4. Clear provenance. The validation team needs to know where the estimate came from, how it was produced, and whether it is independent of the bank’s own modelling process.

No single source necessarily satisfies all four criteria, which is why validation teams tend to combine them. Agency ratings bring depth of analysis on the names they cover, financial-statement models bring breadth and consistency, and peer-bank estimates can provide a more like-for-like comparison with an IRB PD. Peer-bank data can also extend coverage into unrated and private entities where other external sources are limited, which is why banks use it alongside those sources.

The next section looks at what that last category offers to a low-default portfolio, and where it fits alongside the others.

The role of peer-bank IRB estimates in external benchmarking

The internal PD under review comes from the bank’s IRB framework. Peer-bank estimates are produced through other banks’ internal models under similar supervisory expectations, so they are more directly comparable. That means less translation is needed than when benchmarking an IRB PD against an agency rating or a market-driven model, because if the benchmark also comes from IRB models, it has to be genuinely independent.

The point is that a benchmark does not need to come from a different regulatory framework to be independent. It needs to be produced outside the bank’s own modelling process. If several other banks have independently assessed the same counterparty, their estimates provide an external view of where the bank’s PD sits. Aggregating those estimates also reduces reliance on the assumptions or judgment of any single contributing institution. 

A concrete example comes from the Bank of England’s July 2026 Financial Stability Report. which demonstrates the value of bank-produced PD estimates as entity-level credit-risk information. Regarding the benchmark source:,for a difficult-to-assess population, the bank used PD estimates produced in banks’ internal rating models as an external source of entity-level credit risk information.

This does not make the peer consensus a “correct” PD. Banks can legitimately reach different estimates because their information, models, and assumptions differ. The value of the comparison is that it gives the validation team a relevant external reference against which to identify and investigate those differences.

For low-default portfolios, where internal observations provide little challenge on their own, like-for-like external comparison can fill an important part of the evidence gap.

Using Credit Benchmark Consensus data for low-default exposures

Credit Benchmark turns peer-bank estimates into a usable external reference point by aggregating credit-risk estimates contributed by 40+ leading global banks, covering 125,000+ public and private entities, 90%+ of which are unrated.

For each borrower, the consensus is accompanied by two useful pieces of context: how many institutions contributed and how widely their estimates differ.

Take a fund counterparty with an internal PD of 0.18% and a peer consensus of 0.31%. Contributor count shows how broad the peer reference is: twenty contributing bank estimates provide a broader basis for comparison than four. Dispersion shows how much agreement sits behind the consensus. If most estimates cluster around 0.31%, the bank’s 0.18% PD would warrant closer investigation. If peer estimates range from 0.10% to 0.60%, the internal PD sits within the observed range, and the difference is less unusual. 

Consensus PD provides the ability to see how broad the peer reference is and how strongly those estimates converge. 

Credit Benchmark Consensus data includes private and unrated entities, extending the comparison to parts of the portfolio where agency coverage may be limited or financial-statement models may not have the required inputs.

For banks that want peer benchmarking incorporated into the IRB calibration process rather than run as a separate periodic exercise, Credit Benchmark and Oliver Wyman developed IRB Nexus, which supports IRB model development, calibration, and validation, particularly for low- and no-default portfolios.

When estimates and the consensus disagree

A peer benchmark is easier to defend when governance around divergence is defined before the comparison runs. Established validation frameworks typically specify when benchmarking takes place, what degree of difference triggers review, and how the conclusion is documented. That prevents the response from being shaped by the result after the fact.

Once a difference is identified, three things should be considered to accurately interpret the result:

  • Direction: Is the bank more conservative or more optimistic than peers? A higher internal PD may reflect deliberate conservatism; a lower PD raises the question of whether the bank is understating risk.
  • Scope: Is the difference isolated or systematic? A few outliers may come from borrower-specific information, overrides or data issues. Broad divergence is more likely to raise a calibration question.
  • Cause: What explains the difference? Differences may come from information available to the bank, rating decisions, calibration assumptions or portfolio characteristics. What matters is whether that explanation is supported and can be documented.

what to do when your estimate and the consensus disagree

Divergence therefore starts the investigation; it does not determine the outcome. Where the difference is understood and supported by evidence, there is no reason for the internal PD to move towards consensus simply to achieve alignment. Similarly, broad agreement with independently produced peer estimates supports the bank’s existing assessment.

The comparison shows where the bank’s assessment sits relative to its peers, and the comparison provides a starting point for understanding any differences.

How model validation teams use external consensus PDs

Validation teams at large banks and financial institutions use consensus PDs across several parts of the risk process. For IRB banks, they can support model benchmarking, impairment-model validation, and increasingly ongoing portfolio monitoring between formal review cycles. Two UK banks show what that looks like in practice.

Benchmarking IRB and IFRS 9 models: One UK high-street bank uses Credit Benchmark to benchmark its IRB and IFRS 9 models. For IRB validation, it can compare its own PD with a peer-only consensus that excludes the bank’s own estimate. That gives the validation team a cleaner external comparison, particularly where too few defaults exist for backtesting to provide much evidence. The bank also uses the results in portfolio packs, allowing differences between internal and peer PDs to be reviewed more regularly.

Early-warning signals for sector exposures: Another UK high-street bank incorporates Credit Benchmark data into its early-warning framework, with different thresholds by sector. When the external consensus moves far enough from the bank’s own view, the difference can trigger further review. This gives the team another signal between formal validation cycles, when realized defaults may provide little timely information.

The two use cases address the same low-default problem in different ways: one provides an external benchmark where realized defaults are scarce; the other helps identify emerging differences before the next formal review. Other institutions use the same external comparison for model challenge, calibration, and independent confirmation:

  • Nordic bank: Used consensus data during an ECB model inspection to show that its internal ratings often moved ahead of the peer view, strengthening the evidence behind its model challenge.
  • Global commodities trading house: Benchmarked 23,000 names against consensus to identify where internal PDs diverged and focus calibration review on the exposures that needed it most.
  • Canadian pension fund: Compared 2,000 names with consensus and found its internal assessments slightly more conservative, giving the team independent support for its existing view.

Making low-default PD estimates more defensible

Sparse internal data does not remove the need for evidence. It means backtesting alone may not provide enough evidence to support the PD. And more history, additional conservatism, agency ratings, and vendor models can all add evidence, but none necessarily provides a like-for-like benchmark for the IRB PD under review.

That makes the choice of external benchmark important. A more defensible comparison covers the relevant counterparties, measures a comparable form of credit risk, and has provenance the validation team can explain.

Peer-bank consensus PDs add that reference point. Because the underlying estimates originate from banks’ internal credit assessment and rating processes, the comparison is more like-for-like. Aggregating views from multiple institutions also gives the validator a broader peer reference than any single external estimate.

Credit Benchmark provides that comparison using anonymized credit views from more than 40 contributing financial institutions, with coverage extending to private and unrated entities, where comparable external evidence can be harder to find. It does not replace the internal model or determine the correct PD. It shows where the bank’s estimate sits relative to independent peer views.

The first step is a coverage check.

Request a Credit Benchmark coverage check to identify which counterparties have consensus PDs available for external model benchmarking.

Request a coverage assessment →

Want to see Credit Benchmark in action?

Schedule a short 30 minute demo and let our team walk you through the platform, demonstrate key capabilities, and answer any questions live.