ADS.finance

Gini, KS and AUC: Grading Credit Models

ADS Team

Author

September 19, 2026

5 days ago

36

views

Share:

In short: Lenders grade credit models on discrimination - how well the model separates accounts that default from those that do not. Gini, KS and AUC are three related ways of measuring that separation. None of them measure whether the predicted probability is correct, which is a different property called calibration.

Key takeaways

  • AUC of 0.5 is random; 1.0 is perfect separation.
  • Gini = 2 x AUC - 1, so they carry the same information.
  • KS measures the largest gap between the good and bad distributions.
  • Discrimination and calibration are different - a model can rank well and still be miscalibrated.

What do the three metrics measure?

MetricWhat it measuresRange
AUC (area under the ROC curve)Probability a random bad scores worse than a random good0.5 (random) to 1.0 (perfect)
Gini coefficientThe same information, rescaled0 to 1
KS statisticLargest gap between cumulative good and bad distributions0 to 1

Gini and AUC are directly related: Gini = 2 x AUC - 1. An AUC of 0.75 is a Gini of 0.50. They are not independent checks on each other; reporting both is convention rather than additional information.

KS is a different construction. It looks for the score at which the separation between good and bad accounts is widest, which is informative because that score is often near a sensible cut-off.

What counts as a good score?

It depends entirely on the portfolio, which is why cross-portfolio comparisons are usually meaningless. A model predicting default on a broad consumer population has more signal to work with than one predicting default on a pre-screened prime mortgage book where almost nobody defaults.

The useful comparisons are internal:

  • Against the model it replaced - did the rebuild actually improve discrimination?
  • Against its own development sample - a large drop on the validation sample indicates overfitting.
  • Over time - deteriorating discrimination signals population drift and that a rebuild is due.
  • Across segments - a model that discriminates well overall may perform poorly on a subgroup, which is both a performance and a fairness concern.

A model that discriminates beautifully on the data it was built from and poorly on a holdout has learned the sample rather than the relationship. That is why validation on data the model has never seen is not optional.

Why is calibration a separate question?

Because ranking and magnitude are different properties. A model can rank borrowers perfectly - every account that defaults scores worse than every account that does not - while systematically predicting a 2% default rate for a population that actually defaults at 5%.

Discrimination metrics would call that model excellent. It would still be wrong in every way that matters for pricing and provisioning, because expected loss calculations use the predicted probability, not the rank.

So lenders measure both. Calibration is assessed by comparing predicted default rates against actual outcomes within score bands, and a model that ranks well but predicts the wrong levels needs recalibration rather than rebuilding.

This matters beyond model teams. Under AASB 9, provisions depend on expected credit losses, which depend on calibrated probabilities. A model that only ranks correctly cannot support the accounting.

Frequently asked questions

What is a good AUC for a credit model?

It depends heavily on the portfolio - a model on a broad population has more signal than one on a pre-screened prime book. The meaningful comparisons are against the previous model, against the validation sample, and over time.

What is the difference between Gini and AUC?

They carry the same information on different scales: Gini = 2 x AUC - 1. An AUC of 0.75 corresponds to a Gini of 0.50. Reporting both adds convention, not insight.

What is the KS statistic?

The largest gap between the cumulative distributions of good and bad accounts across the score range. It is useful because the score at which the gap is widest often indicates a sensible cut-off.

What is calibration?

Whether the predicted probabilities match observed outcomes - a 2% predicted default rate should produce about 2% actual defaults. It is distinct from discrimination, and it is what matters for pricing and for expected credit loss provisioning.

Related reading

Sources

  • Prudential Standard APS 113 Capital Adequacy: Internal Ratings-based Approach — APRA
  • AASB 9 Financial Instruments — Australian Accounting Standards Board

Information current as at 2 September 2026.

General advice warning: This article contains general information only. It does not take into account your objectives, financial situation or needs, and it is not personal credit or financial advice. Consider whether it is appropriate for you and seek advice from a licensed credit representative before acting.

Any interest rate shown is an example only and is not an offer of credit. Where a rate is quoted, the applicable comparison rate is available from the relevant lender and should be considered alongside it.

Need Financial Assistance?

Connect with our network of trusted finance providers to find the right loan solution for your needs.