Jun 9, 2026arXiv:2606.10333

Privacy-Preserving Credit Risk Prediction with Alternative Data

Hongzhe Zhang, Jiarong Xu, Jing He, Xiao Fang

AI Summary

This paper addresses the challenge of credit risk prediction by introducing PrivacyCredit, a novel machine learning method that enables the secure integration of alternative data while preserving consumer privacy. The study highlights the importance of maintaining model confidentiality and achieving lossless performance, demonstrating that PrivacyCredit can match the predictive accuracy of traditional methods without compromising privacy. Extensive experiments reveal that this approach not only protects sensitive information but also enhances the robustness of credit assessments by leveraging alternative data sources.

Key Contribution

PrivacyCredit achieves top-tier credit risk prediction accuracy without compromising consumer privacy or model confidentiality, setting a new standard for secure financial assessments.

Abstract

Credit risk prediction is a critical problem in the consumer credit industry. Traditionally, financial institutions construct credit risk prediction models using borrowers' demographic, financial, and credit history data, collectively referred to as traditional data. Recent studies have demonstrated that alternative data, such as borrowers' mobile phone communication data, enable lenders to acquire fuller and more accurate profiles of borrowers' creditworthiness, thereby improving credit risk prediction performance. Nevertheless, alternative data are held by external entities independent of financial institutions. Directly sharing alternative data with financial institutions infringe on consumer privacy, yet existing credit risk prediction studies largely overlook this issue. To address this gap, we define a new problem, namely privacy-preserving credit risk prediction with alternative data, which simultaneously considers three practical constraints: the privacy-preserving constraint that protects consumer privacy, the model-confidentiality constraint that learns and stores the model centrally at the financial institution, and the lossless constraint that maintains the performance of the learned model. To solve this problem, we develop PrivacyCredit, a novel privacy-preserving machine learning method. We then theoretically demonstrate the privacy-preserving, model-confidential, and lossless properties of PrivacyCredit. Through extensive experiments using a real-world credit dataset linked with alternative data, we demonstrate the predictive value of securely incorporating alternative data into credit risk prediction and show that PrivacyCredit achieves the same predictive performance as the model learned from the insecure plaintext combination of traditional and alternative data. We further evaluate its model-confidentiality property and computational efficiency.

Data Curation & Synthetic Data Recommendation & Information Retrieval

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

Privacy-Preserving Credit Risk Prediction with Alternative Data

Related Papers