IFRS 9 Implementation Guide

IFRS 9 Implementation Guide: Solving the PD Input Problem for Unrated Borrowers

Share
Repost

The biggest challenge credit practitioners face when operationalizing IFRS 9 into an auditable ECL model is typically the input that carries the most weight: PDs. It drives the ECL calculation and draws scrutiny for three key reasons:

  • Circularity. The standard route starts with IRB estimates, then adjusts and validates them internally. Though technically sound, it still fails an auditor’s test because every step traces back to the bank’s own opinion.
  • No external check available. The usual outside anchors (agency ratings, market prices) don’t exist for private, unrated borrowers, so there’s nothing independent to validate against. It’s why 81% of supervisors say data availability and low data quality are the biggest challenges with setting up ECL models. 
  • Too few defaults. For the portfolios under the most scrutiny (large corporates, financial institutions, private mid-market borrowers), defaults are naturally rare. And 2020 and 2021, which could have added valuable data, were skewed by government support programs that artificially suppressed defaults. 

Two equally defensible PD methodologies can put impairment on the same exposure anywhere between $0.5M and $2.5M, which is essentially a five-fold swing on a single name. Scale that across a book, and the PD methodology an institution chooses effectively decides the provision charge, and the provision charge decides what the earnings release says.

This article covers the biggest issues auditors are flagging, provides a deeper explanation of why PDs attract scrutiny, and outlines options for point-in-time PDs that can be defended when the borrower is private and unrated.

 

IFRS 9 ECL Model Weaknesses

Across multiple research reports, supervisors and auditors keep independently flagging the same symptoms in the numbers. They sometimes don’t name circularity or data scarcity directly, but the results of those three issues show up in many ways:

  • PDs that vary widely across comparable portfolios
  • Staging decisions that lag reality
  • Models that don’t respond to changing forecasts
  • A heavy reliance on overlays

In March 2026, the IMF published a technical note on IFRS 9 from a supervisor’s perspective, implying that there’s too much room for self-serving judgment in ECL estimates, provisions that understate real losses, overlays that need tighter governance, and capital positions that look healthier than the underlying credit risk.

The EBA further proves this is a problem across banks. Its benchmarking of EU banks shows significant variation in 12-month PDs across similar portfolios, and supervisors often can’t tell how much reflects real risk differences versus modeling choices. The same reviews found lenient SICR (Significant Increase in Credit Risk) thresholds delaying stage transfers, and PDs that barely move even when the economic outlook shifts (ironic, since IFRS 9’s whole premise is forward-looking sensitivity).

Furthermore, the PRA’s September 2025 letter raised concerns about elevated model risk, credit drivers that don’t align with what the models were built to detect, overlays requiring more scrutiny, and historical bias in LGD recovery assumptions. And these concerns have remained largely the same over the seven years since the letter was published.

None of these reports shows confusion about the framework. Banks understand the three stages and what IFRS 9 requires. Rather, the slow staging, unresponsive models, and overlay dependence all trace back to the inputs feeding the framework, above all, one input in particular.

 

Why Auditors and Supervisors Challenge PD Estimates More Than Any Other ECL Input


PD estimates drive both the loss calculation and the staging decision

Of the three figures that form the basis of ECL models, only PD does a second job. It’s the trigger for staging, as transfers into Stage 2 are based on the comparison between the current PD and the one at origination, in line with guidance from the Global Public Policy Committee (GPPC).

Qualitative signals still matter as a backstop. Watchlist status or forbearance can flag deterioration before a model catches it, and any loan more than 30 days past due is automatically presumed to have deteriorated, regardless of what the model shows. But these are supplementary checks, not the primary driver. Within the framework described by the GPPC, the PD movement is what actually determines staging.

This means PD does two jobs at once: it sizes the loss, and it decides which loss gets sized (12-month or lifetime). Get it wrong, and both the number and the bucket it lands in are wrong. No other input has that kind of reach, which is why the supervisory findings above, despite using different language, all trace back to the same root cause.

 

IFRS 9 demands a kind of PD that IRB models were never built to produce

IFRS 9 requires four properties at once, none of which IRB models were built to deliver by default:

  • Point-in-time. The PD must reflect conditions right now, at the reporting date, not the smoothed, through-the-cycle average IRB models are designed to produce.
  • Full lifetime coverage. It needs a full term structure, whereas regulatory PD modeling stops at 1 year.
  • Forward-conditioned. The number has to genuinely move when macroeconomic scenarios change. The PDs that the EBA found to be barely responsive to GDP forecasts fail this requirement directly.
  • Unbiased. In GPPC’s own words, there must be no margins of conservatism, no regulatory floors, no supervisory add-ons. A PD that’s deliberately overstated to be “safe” is just as wrong for accounting purposes as one that’s understated.

None of these four properties is optional. And none of them come built into the IRB models most institutions already have running.

 

Converted PDs rest on a chain of discretionary calls

Since almost no institution builds its IFRS 9 PD from scratch, the standard approach is to strip out the built-in conservatism of an existing IRB estimate and expand it into a lifetime curve. Every step of the process involves discretionary choices, such as which method to use to fit the curve, which method to use to convert the one-year IRB PD to a lifetime PD, and how heavily each economic scenario is weighted. 

Institutions starting from similar IRB estimates, therefore, arrive at significantly different IFRS 9 PDs, which is exactly where the $0.5M to $2.5M impairment swing mentioned earlier comes from. The effect widens along the curve: two extension methods that sit close together at year one can be far apart by year five, and a lifetime calculation counts every year of that divergence. And since Stage 2 loans are provisioned off the full curve, that accumulated difference flows straight into the loss estimate for every Stage 2 exposure.

same one-year pd two lifetime curves

The default data behind these estimates is thin

PDs are built from observed defaults, and for the portfolios that draw the most audit scrutiny, defaults are rare. That means the entire term structure often rests on just a handful of historical events. 2020 and 2021 make this worse. 

These two years could have supplied useful new default data, but government support programs artificially suppressed defaults during that period, so the data from those years doesn’t reflect actual conditions. Banks have handled this in three ways: including those years, excluding them entirely, or adjusting for them. Each choice shifts the resulting curve differently. 

Institutions haven’t ignored this problem. They’ve tried the obvious ways to improve default data, but each approach runs into its own limitations.

 

The Three Standard Routes to a Defensible PD

Converting IRB PDs to point-in-time estimates

Most institutions convert the IRB models they already have, which are approved for regulatory capital and have supporting data infrastructure. However, it results in the issue already mentioned in the opening of this piece: circularity. 

Regardless of how carefully conservatism is stripped and calibration shifted, the output still rests entirely on the institution’s own model and default history, and the validation that signs it off is performed by the institution’s own second line. This provenance issue becomes a challenge when an auditor requests independent evidence supporting the estimate.

 

Relying on agency ratings and market-implied PDs

If internal evidence can’t be independent, the logical next move is to look outside the bank. For public entities, a rated counterparty’s agency grade, or a PD derived from bond or CDS pricing, serves as an external opinion. But the limitation is coverage.

Agency ratings mostly cover large public companies, often at the parent level, while the actual exposure is with a subsidiary. Market-implied measures only work for companies with publicly traded stocks, bonds, or CDSs, which excludes most mid-market and private entities. They are also risk-neutral by design (so they overstate real-world default probability) and prone to spiking on sentiment during stress, exactly when a stable estimate matters most. This essentially means external anchors exist only for exposures that need them the least.

 

Compensating with overlays and post-model adjustments

Post-model adjustments became the industry’s standard fix for risks the models miss, and during the pandemic, they were genuinely necessary. The problem is that they have since become permanent fixtures rather than stopgaps, and supervisors flagged it as a major flaw. 

The ECB’s 2024 review of ECL overlays found roughly a third of banks over-relying on subjective judgement, broad umbrella overlays applied across risk categories without differentiation, and lifetime losses assigned to vulnerable sectors without measuring the actual impact or the stage transfers that should follow. 

Banks arguing that their legacy macro-overlays would catch novel risks were reminded that models built before 2018 cannot, by definition, detect risks that did not exist when they were designed. Adding to that is the fact that an overlay, however rigorous, doesn’t produce external validation, which auditors request for.

What’s evident across all three approaches is the absence of an external reference that covers private entities without public ratings, the importance of which is further strengthened by guidance from auditing bodies. 

Before IFRS 9 went live, the six largest audit networks jointly published an implementation paper recommending external ratings and benchmarking for cases with thin default history. Although the paper was published in 2016, the IMF’s 2026 technical note still cites it. So, the expectation that internal data should be checked against external evidence remains relevant today.

What that paper never answered is the question this leaves open: for a portfolio where most borrowers have no rating at all, what external reference is actually available?

 

Consensus Data: The External Reference That Exists Where Ratings Don’t

There’s a real irony in saying unrated borrowers have no external opinion, since they’re actually assessed constantly. Every bank that lends to a private company forms its own credit view of it, rates it internally, and keeps that rating up to date using models that a supervisor has already reviewed and approved. So a mid-market borrower with five lenders has five current, independently produced, regulator-governed credit opinions. 

The problem is then not a lack of assessment; it’s a lack of access. Each opinion is locked inside the bank that produced it, invisible to the other four lenders and to anyone else. 

Consensus credit data solves this by pooling those views. Banks contribute their internal credit estimates, the estimates get combined and anonymized, and the result is a single consensus measure of default risk for each borrower built from the banks that hold actual exposure to it. This directly addresses all three problems described above:

  • Provenance. The input isn’t one bank’s model; it’s many banks’ models, each built and validated independently under its own supervisor.
  • Coverage. Because contributing banks rate the companies they actually lend to (not just companies that issue public debt), the pooled data reaches exactly the private, mid-market, and subsidiary borrowers that agency ratings and market prices miss.
  • Independence. For any single bank using this data, the consensus is genuinely external, being an estimate it didn’t produce, and one it can test its own view against.

This is external benchmarking in exactly the sense the auditors advocate, as it is applied only to the segment lacking an external rating. It doesn’t replace other sources. Agency ratings remain the reference for the issuers they cover, and internal models continue to drive every ECL calculation. Consensus data occupies the space between them: the first external reference for names that never had one, and the second for names that did.

 

Where Credit Benchmark Fits

Credit Benchmark provides this kind of consensus data. It aggregates internal credit estimates contributed by more than 40 global banks into consensus ratings and PDs for around 125,000 legal entities, roughly 90% of which have no traditional agency rating. The data refreshes weekly to twice monthly, depending on the dataset.

In an IFRS 9 context, this data is used in three ways, each addressing one of the problems covered above.

First is benchmarking PD inputs. The consensus provides model and validation teams with a peer comparison to check whether their internal PD for a given borrower aligns with the collective view of other banks lending to that borrower, and to justify ratings.

A bank uses consensus data to pressure-test and refine the PDs in its internal models, bringing them closer to market and peer views. As its Chief Credit Officer put it: “The ability to benchmark our internal views has significantly increased our confidence in risk decisions, and it’s directly informed model recalibration efforts.”

From there arises another application: coverage of unrated and private borrowers. The dataset covers the borrowers that agencies and market prices ignore, down to the specific legal entity actually holding the debt, rather than at the parent company level. That’s why State Street uses it to benchmark internal ratings against market consensus for newer types of counterparties, where internal modelling assumptions require external validation.

Lastly, the data’s frequent updates and structure can support (though not replace) a SICR assessment, as the direction of consensus, upgrade and downgrade activity, and the extent to which contributing banks agree or disagree on a name can all inform that judgment. Sustained improvement in the consensus, for example, is the kind of evidence that supports the release of a Stage 2 provision. This is supporting evidence, not a staging methodology, and it supports the institution’s staging decision.

However, it’s worth noting that Credit Benchmark doesn’t produce ECL numbers, nor is it a turnkey IFRS 9 solution. That’s still the institution’s model; it’s simply one input into a validation and audit process the institution still owns. Additionally, it isn’t automatically “regulator-approved” evidence. Auditors still expect any institution relying on external benchmarking to justify its approach and document its limitations.

 

Closing the PD Gap

The three reasons PDs draw scrutiny (circularity, no external check, and too few defaults) all point to the same solution:

  • Circularity is solved by testing internal PD estimates against a reference the bank didn’t produce itself.
  • Consensus data covers the private and unrated borrowers that agencies and markets never reached, serving as a reliable external check.
  • The too-few-defaults problem eases because the evidence base is no longer one bank’s thin history. Rather, it’s the combined, regulator-supervised views of many banks holding the same exposures.

But having this data available isn’t enough; how it’s used matters just as much, as supervision is shifting toward comparison. The IMF describes regulators now building their own “challenger” models and requiring regular benchmarking. Reviews are moving away from “show us your framework” and toward “explain why your numbers differ from everyone else’s.”

This shift changes what “defensible” means in practice. A benchmark pulled together once, in response to a single audit finding, answers that one review and then goes stale. Banks in the strongest position are the ones already consistently checking internal PDs against consensus data as part of normal model monitoring, so that staging decisions and provision releases are backed by more than just the bank’s own judgment.

A methodology choice can still swing impairment on a single name by up to five times, and the resulting provision charge still determines what the earnings release says. However, now, they can get more defensible.

Schedule a demo

Please complete the form below to arrange a demo.

    By submitting this form you agree to Credit Benchmark’s
    Privacy Policy and Terms and Conditions.

    Subscribe to our newsletter

      By submitting this form you agree to Credit Benchmark’s
      Privacy Policy and Terms and Conditions.