CERESResearch Repository

Default Prediction of Unlisted Firms

Loading...
Thumbnail Image

Date published

Free to read from

2026-06-19

Supervisor/s

Industry supervisor/s

Journal Title

Journal ISSN

Volume Title

Department

BAM

Type

ISSN

Format

Citation

Abstract

This study addresses the critical gap in default prediction for unlisted companies, which constitute over 99% of global enterprises yet remain systematically understudied compared to listed firms. Using the FAME database encompassing over 60,000 UK private companies from 2010-2024, the research systematically compares four machine learning approaches: logistic regression, XGBoost, LightGBM, and neural networks for bankruptcy prediction in extremely imbalanced datasets. The methodology employs principal component analysis to identify seven core financial variables (ROA, AssetTurnover, Profit margin, Cash ratio, Current ratio, CashInterestCover, and Company age) spanning profitability, operational efficiency, and solvency dimensions. Time series cross-validation ensures robust evaluation across different economic cycles, including the COVID-19 pandemic period. Key findings reveal XGBoost's superior performance, achieving 91.52% recall rate by successfully identifying 546 of 597 bankrupt companies whilst missing only 51. This significantly outperforms logistic regression (84.77%), LightGBM (67%), and neural networks (12.4%). Critically, the study demonstrates that traditional ROC-AUC metrics prove misleading under extreme class imbalance (0.35%-0.91% default rates), with all models achieving 0.74-0.93 AUC values primarily through correct non-default classification rather than genuine bankruptcy identification capability. PR-AUC emerges as a more reliable evaluation metric. The research provides practical guidance for financial institutions, recommending XGBoost for maximum risk coverage and logistic regression for regulatory compliance scenarios. Neural networks demonstrate complete failure, highlighting deep learning limitations in extremely imbalanced sample environments. These findings contribute significantly to financial risk management theory and practice, offering evidence-based algorithm selection frameworks for credit risk assessment in the private company sector.

Description

Software description

Software language

Git repository

Keywords

default prediction, unlisted companies, machine learning, XGBoost, logistic regression, neural networks, LightGBM, imbalanced data, ensemble learning

DOI

Rights

Funder/s

Grant number

Relationships

Relationships

Resources