Credit Risk Modeling

Credit Risk Modeling: A Practitioner’s Guide to Models, Validation, and Regulatory Frameworks

Share
Repost

Credit risk modeling is a discipline where the cost of getting it wrong shows up in failures that can destabilize entire financial systems. We don’t need to go back too far to get an example.

In 2008, models across the financial system failed together. They had been calibrated during a period of stable housing prices and low defaults. So when conditions shifted, the models had no reference point for what was actually happening. Default correlations turned out to be far higher than assumed, and recovery rates collapsed alongside the spike in defaults.

The regulatory response was designed to prevent those failures from repeating. But there’s still a large gap: credit risk is increasingly hard to model, because commercial portfolios today are increasingly exposed to private companies, unrated counterparties, and cross-border entities where that data is thin.

Below, we explore credit modeling for entity-level portfolios, including the regulatory framework, the model types, and the lifecycle from development through recalibration. We also explain how to deal with the data scarcity issue, by giving a practical example of how other financial institutions are tackling it.

 

How Model Outputs Shape Capital, Provisions, and Credit Decisions

A credit risk model takes economic conditions and entity-specific inputs (financial ratios, leverage, industry outlook, and management quality). Then, converts them into a measurable output, which is most often a probability of default or a credit spread. This converts credit judgment into something quantifiable, comparable across a portfolio of hundreds or thousands of entities, and defensible when a regulator or auditor asks how a decision was reached.

The people who build these models, typically quantitative analysts, IRB teams, and model risk functions, are not the same people who use the outputs:

  • Credit committees rely on them to approve or decline exposures.
  • Relationship managers use them to price risk. 
  • Capital planning uses them to allocate reserves across the book. 
  • A credit officer cares about whether the PD on a specific counterparty is right.
  • Regulatory supervisors use them to assess whether the institution holds enough capital. 

Each of these consumers asks a different question of the same number, and each has a different tolerance for error. 

The three parameters that credit risk models are built to estimate, whether individually or in combination, are captured in the expected loss formula:

EL = PD × LGD × EAD. 

Economic capital graph

This is why model accuracy is key. Under Basel IRB, the regulatory capital charge is sized against unexpected loss. If the PD feeding the formula is too low, the bank holds less capital than its actual risk requires. If it’s too high, capital is locked up that could be deployed elsewhere. Either way, the error significantly affects how much capital the institution holds against its book.

 

The Regulatory Framework Guiding Credit Risk Modeling

Every credit risk model exists inside a regulatory structure that determines how its outputs are used. Basel determines how PD, LGD, and EAD estimates translate into the capital a bank must hold. SR 26-2 defines what evidence supervisors expect when they ask whether the model is working as intended. IFRS 9 and CECL dictate how those same model outputs flow into provisioning and hit the income statement.

 

Basel: How Model Outputs Become Capital

The Basel framework is the international standard that determines how much capital banks must hold against their credit exposures. For credit risk modelers, Basel matters because it determines how much capital an institution holds to hedge against risk. The framework has evolved over four iterations, each giving banks freedom to use their own models while tightening the standards those models must meet:

  • Basel I introduced a flat capital rule. Every corporate loan carried the same 100% risk weight regardless of credit quality. A loan to a strong investment-grade company required the same capital as one to a highly leveraged mid-market borrower.
  • Basel II introduced risk sensitivity. The internal ratings-based approach lets banks use their own PD, LGD, and EAD estimates to calculate risk weights. A borrower with a low PD and strong collateral produces a lower risk weight. A borrower with a high PD and weak recovery prospects produces a higher one. The more accurately a bank estimates default risk, the more precisely its capital reflects the actual risk in its portfolio.
  • Basel III raised the bar on capital quality. The minimum CET1 ratio moved to 4.5%, with conservation and countercyclical buffers on top. But the way risk weights are calculated stays largely the same.

Basel IV is the current iteration, and the one that connects most directly to modeling decisions today.

Under previous versions of Basel, banks with strong internal models could use them to significantly reduce their capital requirements. Basel IV limits that by introducing an output floor that says, no matter how sophisticated a bank’s internal model is, the resulting risk weights cannot fall below 72.5% of what the simpler standardized approach would produce. In other words, internal models can still reduce capital, but only up to a point.

This has a specific impact on unrated entities. Under the standardized approach, any entity without an external credit rating automatically carries a 100% risk weight. But if a bank can support a lower risk classification for that entity, through a credible external rating or consensus-based assessment, the capital it must hold drops accordingly. For example, moving a $1B unrated exposure from a 100% to a 65% risk weight frees up approximately $2M in capital.

For modeling teams, before Basel IV, recalibrating a model was primarily a documentation exercise. Now, a recalibration that lowers PDs directly reduces the capital the bank holds against those exposures. Supervisors will ask what evidence supports the change, particularly for unrated entities where internal data alone may not be sufficient. Later on, we’ll explain where some financial institutions get this evidence.

 

SR 26-2: Model Risk Management

SR 26-2 covers how models are developed, validated, and governed, and serves as the reference point examiners use when assessing whether an institution’s credit risk models are fit for purpose.

Technically, SR 26-2 sets expectations for managing risk, rather than enforceable legal requirements. In practice, however, examiners evaluate models against SR 26-2’s framework, and institutions that fall short receive formal findings with remediation timelines. The guidance is expected to be most relevant to banking organizations with over $30 billion in total assets.

SR 26-2 organizes model validation around three elements:

  • Conceptual soundness. Does the model’s theory hold up? Are its limitations documented? Examiners look for evidence that the development team understood why the model works, not just that it passed statistical tests.
  • Ongoing monitoring, including benchmarking. Is the model’s performance tracked over time? Is it compared against external reference points? SR 26-2 describes external benchmarking as a key element of effective validation. Most examiners now expect to see it during reviews, particularly for portfolios where internal data alone is too thin to support independent validation.
  • Outcomes analysis. Do the model’s predictions match what actually happened? This covers backtesting against realized defaults, calibration checks across rating grades, and assessments of how well the model distinguishes defaulters from non-defaulters.

When gaps are found in any of these three areas, they surface as Matters Requiring Attention. These are formal examiner findings, each with a remediation timeline. Unresolved MRAs can lead to capital add-ons, restrictions on model use, or limits on business activities until the issue is fixed.

 

IFRS 9 and CECL: How Model Outputs Flow Into Provisions

Rather than waiting for a borrower to default before booking a loss, IFRS 9 requires banks to estimate expected credit losses on a forward-looking basis from the moment a loan is originated. It does this through a three-stage impairment model:

  • Stage 1: The exposure is performing normally. The bank provisions for 12 months of expected losses.
  • Stage 2: Credit risk has increased significantly since origination. Provisions now cover the full remaining life of the exposure, which is a substantial step up.
  • Stage 3: The exposure is credit-impaired. Provisions cover lifetime expected losses, with additional recognition of incurred losses.

Staging decisions depend on point-in-time PDs. Through-the-cycle ratings are designed to be stable across the economic cycle, which means they won’t flag deterioration until it’s severe. PIT PDs reflect current conditions and are sensitive enough to catch the early shifts that staging is designed to detect.

Moving an exposure into Stage 2 is straightforward; the harder part is justifying the move back. Auditors will ask for evidence that credit risk has genuinely improved, based on factors such as sustained financial improvement, a rating upgrade, or external benchmarks confirming that peer assessments have also improved. Where that evidence does not exist, the migration is difficult to defend.

CECL is the US equivalent of IFRS 9, applicable to US GAAP-reporting institutions. Instead of three stages, CECL requires banks to estimate lifetime expected losses from the point of origination for all exposures, regardless of whether credit risk has changed. The structure is simpler, but the underlying challenges are the same: building PD curves across the full maturity of each exposure, producing forward-looking assumptions that withstand audit scrutiny, and covering private and unrated entities where historical loss data is limited.

It is worth noting that IFRS 9 and SR 26-2 operate in different domains. IFRS 9 governs how losses are recognized in financial statements. SR 26-2 governs how the models producing those estimates are developed and validated. In practice, the two overlap as a model used for IFRS 9 staging will also fall under SR 26-2 governance, but they impose different obligations and should be treated as such in documentation.

With the regulatory context established (what capital requires, what supervisors expect from validation, and what accounting standards demand of provisioning), the question becomes: what types of models do institutions use to meet these requirements, and how does the choice of model depend on the data available?

 

The Main Model Types and the Tradeoffs Between Them

Banks don’t use a single credit risk model. Structural models may serve public corporations where market data exists, statistical models for mid-market and private entity portfolios, reduced-form models for derivatives pricing and CVA, and portfolio models for capital allocation and stress testing.

Three things drive which model applies where: the data available for that segment of the portfolio, the output the model feeds into (capital, pricing, provisions, or derivatives valuation), and which regulatory framework governs that output. 

This is why credit risk models are best understood along two dimensions. The first, by what’s being modeled, explains why a bank needs multiple models: PD, LGD, and EAD each require separate estimation, and portfolio loss requires another approach entirely. The second, by modeling approach, explains why the same parameter is estimated differently for a publicly traded corporation than for a private mid-market borrower.

 

Models by Risk Component: PD, LGD, EAD, and Portfolio Loss

 

Probability of Default (PD) Credit Risk Model

PD is the estimated likelihood that an entity defaults within a given time horizon. For Basel capital calculations, that horizon is one year. For IFRS 9 and stress testing, banks need multi-year PD curves extending across the full remaining maturity of each exposure.

Calibrating a PD model means fitting observed default frequencies to rating grades. And the relationship between rating grade and default rate is exponential. This means a one-notch downgrade from AA to a might move the default rate from 0.02% to 0.05%. Meanwhile, a one-notch downgrade from b to ccc raises the risk premium from 3.5% to 15%.

This is where entity-level modeling runs into its core data problem. For private companies, unrated counterparties, and low-default portfolios, there are often too few observed defaults to fit a reliable calibration. A rating grade that has produced zero defaults in a five-year window hasn’t been validated by the absence of defaults. It simply hasn’t been tested. This can be solved by either:

  • Supplementing with external benchmarks 
  • Applying conservative floors to low-default grades

In either case, the justification for each choice must be documented.

The other source of complexity to consider is the development of multi-year PD, which IFRS 9 recommends. 

Entities migrate between rating grades over time, and a bbb-rated borrower that drifts to bb in year three and b in year four produces a cumulative curve that accelerates rather than running in a straight line. These curves are built either by compounding annual credit transition matrices forward or by fitting survival functions to multi-year default data. 

These curves are typically built either by compounding annual credit transition matrices forward or by fitting survival functions to multi-year default data, both of which are covered in detail in our guide to PD modeling. In either case, validating the resulting term structures against external reference curves, where available, is the most reliable confirmation of the projection.

 

Loss Given Default (LGD) Credit Risk Model

LGD is the proportion of exposure a bank does not recover after a borrower defaults. It’s built by:

  • Segmenting defaulted exposures by collateral type, seniority, and jurisdiction
  • Analyzing historical recoveries on those segmented defaulted exposures
  • Applying downturn adjustments to reflect the stress conditions when most defaults actually occur

Although this might sound simple, in practice, each of those steps introduces complexity.

First, recovery takes time. A secured corporate loan backed by commercial real estate might recover 70 cents on the dollar, but if enforcement takes two years, the present-value recovery is materially lower once legal costs and property depreciation are factored in. Even the choice of discount rate can shift the estimate by several percentage points. 

The practical way to take this into consideration when modeling is to document the choice of discount rate explicitly, apply it consistently across the portfolio, and separate cure rates from loss-event recoveries so the two don’t contaminate each other.

Another layer of complexity arises from the lack of m-n collateral structures, a mechanism for overcollateralization or collateral pooling. It indicates a ratio where the value of the pledged assets (M) exceeds the value of the secured obligation (N).

Unlike retail lending, where collateral mapping is typically one facility to one asset, a single borrower may have multiple facilities with different seniority, secured against a mix of collateral, with some pledges dedicated and others shared. Which collateral covers which exposure, in what order, depends on the intercreditor agreement and the jurisdiction. 

The common fix is to track which collateral covers which loan, model the recovery order by seniority, and estimate LGD separately for each facility rather than applying a single average recovery rate across the entire borrower relationship. That way, you don’t misallocate capital across the entire relationship.

 

Exposure at Default (EAD) Credit Risk Model

EAD is the amount a borrower owes at the moment they default. For a term loan, this is simply the remaining principal. For revolving credit lines and derivatives, it is harder to pin down because the exposure itself moves.

A company with a $50M revolving facility might normally use $20M. But as financial pressure builds, it draws more of its credit line. By the time it defaults, utilization may be close to the full $50M. This pattern is consistent, as borrowers use more of their available credit as their condition worsens. 

To capture this, EAD models include a credit conversion factor (CCF), which is the share of the undrawn balance expected to be drawn before default. However, because a single portfolio-wide estimate will be skewed, modelers create separate conversion factors by rating grade and facility type. That way, the final model reflects the link between deteriorating credit quality and rising utilization.

Derivatives work differently. On a five-year interest rate swap, the bank’s exposure to its counterparty changes over the life of the contract. It starts low in year one, when rate movements are small, and peaks somewhere in the middle, when uncertainty is highest and substantial payments remain. Then, it falls again toward maturity as payments settle. 

Because exposure moves like this, modelers build a profile of expected exposure across the full contract life, using simulation or SA-CCR (the standardized regulatory formula for calculating derivatives exposure).

The expected loss formula assumes PD, LGD, and EAD move independently. In practice, they move together and sometimes overlap. When defaults spike, collateral values fall, linking LGD to PD. When credit quality deteriorates, borrowers use more of their available credit, linking EAD to PD. And when more is drawn, there is more exposure competing for the same recovery proceeds, linking EAD to LGD.

standardized regulatory formula for calculating derivatives exposure

The formula remains the regulatory standard and is accepted as a working simplification, but you should understand that it systematically understates loss under stress, and document that limitation rather than hide it. 

 

Portfolio models

Individual PD, LGD, and EAD estimates describe the risk of a single borrower. Portfolio models ask a different question: what happens when a bank lends to thousands of borrowers at the same time? If those borrowers had nothing in common, their defaults would be spread out and stay close to what the bank expected. 

However, in reality, borrowers are all exposed to and affected by the same economy. When GDP contracts, interest rates spike, or an entire sector comes under pressure, defaults arrive together, rather than one at a time. This means a bank with 500 industrial borrowers in one region will see far more defaults arriving together than a bank with 500 borrowers spread across sectors and geographies, even at the same average PD.

The Basel IRB capital formula accounts for this by incorporating asset correlation into the risk-weight calculation. However, it applies the same correlation assumption whether a portfolio is concentrated in one sector or spread across twenty. Therefore, it’s important to supplement the formula with internal concentration analysis to capture how severe losses can actually get on the most exposed portfolios.

 

Models by Methodology: Structural, Reduced-Form, Statistical, and Machine Learning

The second dimension is how the model works. Four approaches are widely used in entity-level credit risk modeling, each built on a different assumption about what drives default and each requiring different data. Which is suitable comes down solely to which one best fits the available data for a given portfolio segment and the required regulatory output.

Models by Methodology

Structural models

Structural models start from the premise that a company defaults when its asset value falls below its debt. Rooted in the Merton framework, the approach treats equity as a call option on the firm’s assets and measures how far asset value sits from the default threshold (known as distance-to-default).

These models yield forward-looking PDs. They capture changes in leverage and asset volatility through market data before deterioration shows up in financial statements, making them useful for large public corporations, financial institutions with traded securities, and CVA calculations.

The limitation is coverage. Asset value cannot be observed directly and must be inferred from equity prices and volatility, so the model only works for entities with publicly traded equity or liquid CDS. Private companies, unrated subsidiaries, fund structures, and most mid-market borrowers, which form the majority of a typical commercial bank’s wholesale book, are excluded.

 

Reduced-form models

Reduced-form models take a different starting point. Rather than tracing default back to a company’s balance sheet, they model it as an event that can occur at any time with a certain probability, known as the default intensity. That intensity can be estimated from market prices such as CDS or bond spreads, or from historical default data, depending on how the output is used.

When calibrated to market prices, reduced-form models are the natural choice for derivatives pricing and CVA. They are also useful for building multi-year PD curves when historical default observations are too scarce to support a statistical calibration.

One distinction matters here: market-calibrated reduced-form models produce PDs that are systematically higher than those from historically calibrated models. That’s because market prices include a risk premium on top of the actual default probability. 

Therefore, market-calibrated PDs belong in derivatives pricing and CVA, while historically calibrated PDs belong in capital calculations and provisioning. Using one where the other is required is a recurring source of audit findings.

Like structural models, reduced-form models only work where traded instruments exist. For entities without CDS or bonds, the default probability must come from another source, such as internal statistical models or external consensus data.

 

Statistical and econometric models

Statistical models estimate default probability from observable characteristics of the borrower (financial ratios, industry classification, macroeconomic conditions) using historical defaults to identify which characteristics best predict failure. Logistic regression and rating scorecards are the most common approaches in entity-level credit.

The most consequential choice when building a statistical credit risk model is how to set the level of its PD outputs:

  • Point-in-time (PIT): PDs reflect current conditions, rising in downturns and falling in expansions. Best suited for IFRS 9 staging and stress testing, where the goal is to capture deterioration as it happens.
  • Through-the-cycle (TTC): PDs target long-run average default rates and remain stable across the economic cycle. Best suited for Basel capital planning and long-run risk appetite frameworks.

Neither is universally correct. However, it’s important to consider what each one is meant for and document your choice. The two measure different things, and mixing them up (e.g., validating a PIT model against a TTC benchmark or presenting TTC PDs in an IFRS 9 staging process) is a common reason models are flagged during regulatory reviews.

Statistical models have one major vulnerability: they can only reflect the conditions they were trained on. 

A model built on five years of low defaults and stable spreads will backtest well against that same period. But when conditions shift due to a rate shock, a sector downturn, or a liquidity crisis, the model enters an environment it has never seen, and its past accuracy offers no guarantee it will hold.

 

Machine learning models

Machine learning models use algorithms that learn patterns from data rather than following a predefined formula. Methods like gradient boosting, random forests, and neural networks analyze large volumes of borrower data, identify which combinations of characteristics best predict default, and improve their accuracy as more data is fed in.

In entity-level credit, their strongest applications are for tasks where overall accuracy across the portfolio matters more than explaining any single entity’s rating. Tasks like:

  • Portfolio surveillance
  • Early warning systems
  • Spotting unusual shifts in rating distributions

The flip side is that this strength becomes a constraint in regulated lending. Examiners expect credit decisions on individual entities to be explainable and documentable, not delegated to an algorithm that cannot articulate why it is assigned a particular rating. Models that cannot meet this standard will be flagged during regulatory reviews.

As such, most institutions use machine learning models for monitoring and surveillance and interpretable models like logistic regression and scorecards for individual credit decisions and regulatory submissions.

 

The Credit Risk Modeling Lifecycle

A credit risk model goes through data preparation, development, validation, monitoring, stress testing, and recalibration. Each stage is a distinct governance obligation, and the choices made early constrain what is possible later.

credit risk modeling rate (1)

Data preparation and feature engineering

Data preparation is the process of collecting, cleaning, and organizing the inputs a credit risk model will use to estimate default. For entity-level models, those inputs typically include:

  • Financial ratios (leverage, interest coverage, profitability)
  • Industry and sector classification
  • Macroeconomic variables (GDP growth, credit spreads, unemployment)
  • Qualitative assessments, such as management quality and market position (if available). 

Each requires consistent historical collection across the portfolio, and building and maintaining that dataset is often a greater effort than building the model itself.

This stage comes before model selection because the data determines which approaches are viable. For example, a modeling method that needs five years of quarterly financials cannot be used for a portfolio where half the entities report annually, and a third don’t file public statements at all.

 

Model development and calibration

Once the data is prepared, the model is built and calibrated. Two decisions at this stage shape everything downstream.

The first is the PIT versus TTC calibration choice, covered in the statistical models section above. Whatever choice is made here commits the institution to a specific path for IFRS 9 staging, stress testing, and regulatory submissions. Changing it mid-lifecycle requires full redevelopment and revalidation.

The second is the PD time horizon. As covered in the PD section above, one-year PDs satisfy Basel capital, but IFRS 9 and stress testing require multi-year cumulative curves, which introduce additional calibration complexity. More on how these curves are built is covered in detail in our PD modeling guide.

Although this stage is meant mainly for development, documentation deserves serious attention. SR 26-2 calls for comprehensive documentation of assumptions, limitations, and intended use. Documentation written alongside development, capturing why decisions were made as they were made, is far more credible to examiners than a retrospective write-up assembled after the model is in production.

 

Model validation

Validation is the stage examiners scrutinize most closely, because it is where the institution must demonstrate that its models actually work. SR 26-2 frames it around three elements, and examiners evaluate all three. Every model must pass all three to adhere to regulatory guidance:

  • Conceptual soundness: Does the model’s theory hold? Are its limitations documented? Examiners look for evidence that the team understood why the model works, not just that it passed statistical tests.
  • Ongoing monitoring, including benchmarking: Is performance tracked over time? Is the model compared against external reference points? External benchmarking is a key element of effective validation under SR 26-2, and its absence is increasingly difficult to defend during reviews.
  • Outcomes analysis: Do predictions match what actually happened? This is assessed through backtesting against realized defaults, calibration checks across rating grades, and measures of how well the model distinguishes defaulters from non-defaulters.

For the third element (outcomes analysis), examiners check two dimensions, and both must hold. The first is discriminatory power, which checks whether the model correctly ranks entities from least to most risky, as measured by the Gini coefficient. The second is calibration, which checks if the PD levels are accurate.

Both dimensions depend on comparing model predictions against actual defaults. This creates a problem for portfolios where defaults rarely occur, as there simply aren’t enough events to run a meaningful statistical test. For portfolios like these, external benchmarking and expert judgment become the primary way to assess whether the model is working. 

 

Ongoing monitoring

Ongoing monitoring tracks whether the model is still performing as expected between formal validation reviews. Without it, a model can drift out of alignment with the portfolio it serves for months or years before anyone notices.

In practice, monitoring teams track several indicators:

  • The population stability index (PSI) measures whether the characteristics of the borrowers the model is scoring have shifted since the model was built. Values above 0.25 typically trigger a formal review. 
  • Rating distribution monitoring tracks month-to-month changes in portfolio PDs and the frequency with which credit officers override the model’s output. 
  • External benchmarking compares internal PDs against peer assessments to check whether the model is drifting relative to the market.

Before deciding how to respond, the monitoring team needs to determine the cause of the shift. If the portfolio has changed gradually (new sectors, different entity sizes, evolving borrower profiles), that is drift, and the model can be recalibrated on updated data. But if the underlying relationship between the model’s inputs and default outcomes has changed due to new rate environments or a systemic stress event (regime change), the model will need to be rebuilt.

 

Stress testing and recalibration

Stress testing asks what happens to a portfolio’s credit risk under adverse economic conditions. Rather than reassessing each borrower individually, the model applies a macroeconomic scenario to the entire portfolio at once and calculates how it would affect default rates across the portfolio. 

Every rating grade is affected at once. Investment-grade PDs might double, speculative-grade PDs might triple, and portfolios concentrated in a single sector or region are hit harder because their borrowers are all affected by the same conditions simultaneously.

IFRS 9 builds on this by requiring banks to provision for more than one possible future. Instead of using a single best-guess forecast, banks are expected to estimate expected credit losses across multiple scenarios and blend them into a single provision number. Here’s how it works:

  • Estimate credit losses under multiple scenarios: Banks model at least three economic outcomes, including a base case, a pessimistic case, and an optimistic case.
  • Assign probability weights to each scenario: Each scenario is assigned a probability. For example: 60% base, 25% pessimistic, 15% optimistic. 
  • Blend into a single provision: The final provision is the weighted average across all three scenarios. 

The standard does not prescribe how to set these probabilities. A bank that assigns 35% to the pessimistic case will provision materially more than one that assigns 15%, even using identical scenarios. However, auditors will ask for the evidence supporting each probability weight.

Stress scenarios also affect IFRS 9 staging. A pessimistic stress scenario can trigger a move of borrowers from one stage to another, which also affects the capital provision for that borrower.

 

Common Mistakes in Credit Risk Modeling

Most modeling mistakes are the predictable result of working with incomplete data, shifting market conditions, and competing demands from regulators, auditors, and the business. Understanding where the process tends to break down is as useful as understanding the methodology itself.

 

Building on Incomplete Data

Private companies, unrated counterparties, and fund structures often lack audited financial statements, agency ratings, and publicly traded instruments. Banks still need to model their credit risk under Basel IRB, but the outputs for these entities rest on thinner data. Teams, therefore, make the mistake of putting these models into production without a way to verify the results externally. This gap persists until an examiner asks for evidence of external validation, and there is none.

 

Overfitting to One Market Environment

Every credit risk model is trained on a specific historical period and will backtest well against that same period. Strong backtesting results only prove the model worked in the conditions it was built on. When a rate shock, sector downturn, or liquidity event falls outside that window, the model enters an environment it has never seen.

 

Choosing the Model Before Understanding the Data

Teams sometimes debate which algorithm to use before examining whether the data supports any of them for the segment in question. Meanwhile, the method should follow the data, not the other way around. A simple model on clean, well-understood inputs will outperform a sophisticated one built on data nobody has properly examined.

 

Treating a Model Output as a Decision

A credit risk model produces a PD, while a credit decision requires a documented judgment (approve, decline, or modify terms) with reasons an examiner can evaluate. However, teams sometimes make the mistake of treating the model output as the decision itself, rather than as one input that a credit officer must be able to explain and defend. This creates a defensibility gap that stays unnoticed until examiners ask for the reasoning behind a specific rating.

 

Relying on a Single Model

Different models examine the same borrower from different angles. A structural model might flag deterioration through falling equity prices before it shows up in financial statements. A statistical model might miss it because the latest financials haven’t been filed. Running both provides an independent challenge. Running only one provides no signal when that model is wrong. The mistake is treating a single model’s output as settled rather than as one view that should be tested against others.

 

Validating Against the Wrong Benchmark

For low-default portfolios, backtesting cannot produce statistically meaningful results because there are too few defaults to test against. To solve this, it’s very easy to make the mistake of filling the gap with benchmarks that don’t provide genuine independent challenge. 

Additionally, internal peer comparison (comparing one desk’s ratings against another within the same institution) reflects the same assumptions back at the institution. Through-the-cycle benchmarks smooth away the cyclical movements that point-in-time validation is designed to detect. Both give false comfort in exactly the situations where real external benchmarking is most needed.

 

No External Reference for Cross-Institution Divergence

Regulatory exercises consistently find that different banks assign different risk weights to the same hypothetical portfolio. Two banks looking at an identical set of corporate borrowers can produce materially different PDs, not because the borrowers are different, but because the models are. 

In some cases, the difference is due to one bank applying more conservative assumptions than another. In other cases, it reflects miscalibration that neither bank can detect internally, because each model’s outputs are internally consistent. Without comparing PD estimates against an independent external source covering the same entities, a bank has no way to know which category its divergence falls into.

Several of these failure points converge on the same structural gap: internal models, however well-built, cannot generate the external reference point they need to validate themselves. External benchmarking closes this gap.

 

Closing the Validation Gap for Private and Unrated Entities

A model can be conceptually sound, well-documented, and statistically strong on its training data, and still be miscalibrated relative to the market without anyone inside the institution knowing. Detecting that kind of error requires a reference point that the institution didn’t produce itself. But finding one that actually covers the unrated entities where validation pressure is highest is the harder problem.

The two most obvious sources of external reference each have a structural limitation. Starting with agency ratings, they:

  • Cover publicly traded corporates and large issuers well, but excludes most private companies, subsidiaries, and mid-market borrowers.
  • Update quarterly or on an event-driven basis, which is too slow for point-in-time validation or early warning.
  • Follow an issuer-paid model, in which the rated company pays for the rating. This is an independence concern that regulators such as the Securities and Exchange Commission have openly raised.

The story is no different for market-implied measures such as CDS spreads. 

  • Only available for entities with liquid traded instruments.
  • Include the risk premium discussed in the reduced-form models section, meaning they overstate real-world default probability and are unsuitable for capital or provisioning without adjustment.
  • Can spike on sentiment and liquidity in stress conditions rather than reflecting actual changes in credit quality.

Consensus credit data plugs the gaps identified above by taking a different approach. Rather than relying on a single agency’s view or on market-implied signals, it aggregates anonymized PD estimates from over 40 global banks, nearly half of which are Global Systemically Important Banks, and publishes them. 

The dataset covers approximately 120,000 entities, around 90% of which are unrated by traditional credit rating agencies. This includes private companies, unrated counterparties, fund structures, and mid-market borrowers. Three characteristics make the data relevant for modeling teams:

  • The data refreshes weekly, supporting point-in-time monitoring and early warning at a cadence consistent with how credit surveillance actually operates. 
  • Every contributing bank’s internal ratings are produced by regulator-validated IRB models, calibrated to actual loss experience and overseen by prudential supervisors. The dataset inherits that governance. 
  • A three-bank minimum aggregation rule ensures no single contributor’s view is identifiable, which is the confidentiality architecture that allows the dataset to exist at scale.

The accuracy of consensus ratings has been independently tested. A like-for-like comparison against S&P over 2015–2024 found consensus ratings achieved a one-year Gini of 0.88 versus S&P’s 0.91, with comparable results at three-year (0.83 vs. 0.85) and five-year (0.81 vs. 0.82) horizons. This meant it had a comparable rank ordering but with substantially broader entity coverage.

KeyBank embedded consensus data into its credit risk models. In the words of Dale Clayton, Chief Credit Officer, KeyBank:

“Credit Benchmark has given us a clearer lens into how our peers assess credit risk, especially for unrated names. The ability to benchmark our internal views has significantly increased our confidence in risk decisions, and it’s directly informed model recalibration efforts.” 

Consensus data is one component of a robust credit risk modeling framework. It provides an independent reference point that internal processes cannot generate, but does not replace internal models, agency ratings, or the institution’s own credit judgment. It only fills the gap in entity-level PD benchmarking for segments of the portfolio that other external sources do not cover.

See how Credit Benchmark closes the coverage gap for unrated borrowers.

Conclusion

Credit risk modeling is a discipline that integrates model design, regulatory compliance, lifecycle governance, and validation into a single continuous process. Any breakdown in the chain can diminish the entire process. For example, a calibration decision in development determines what IFRS 9 staging can defend, and a data gap in preparation limits what validation can test.

The hardest part of that process today is the data. Portfolios are increasingly exposed to private entities, unrated counterparties, and cross-border borrowers, where the external reference points on which validation depends are thinnest. Internal models can be well-built and well-documented and still lack the independent benchmark needed to confirm they are calibrated against reality. 

Closing that gap through a combination of rigorous internal development and credible external benchmarking is where the discipline is most actively evolving.

See how consensus data fits into your modeling framework

See how Credit Benchmark’s consensus data for unrated entities fits into your modeling framework.

Schedule a demo

Please complete the form below to arrange a demo.

    By submitting this form you agree to Credit Benchmark’s
    Privacy Policy and Terms and Conditions.

    Subscribe to our newsletter

      By submitting this form you agree to Credit Benchmark’s
      Privacy Policy and Terms and Conditions.