HIMANSHU.KUMAR

Data science and ML application

Credit Risk Explorer

An educational app that turns applicant details into an explained credit-risk estimate.

Role
Data science · applied machine learning
Date
Nov 2025
Stack
Python · XGBoost · scikit-learn · Streamlit · pytest
Links
Repository ↗Live demo ↗
Credit Risk Explorer dashboard with clearly labelled applicant and financial inputs beside the model result workspace.
Working Credit Risk Explorer screen showing labelled inputs and a completed, explained educational assessment.

Case in one minute

Problem
Frontend inputs can drift from model training
Solution
Validate 8 fields and reuse one saved pipeline
What I built
Streamlit UI, preprocessing, training, and tests
Result
Explained output backed by 9 behavior tests

01 / Problem

Problem

A model result becomes unreliable when the interface encodes inputs differently from the training pipeline.

The app also needs to reject invalid values before inference and explain the output without presenting an educational score as a lending decision.

02 / Solution

Solution

Validate all eight inputs, then run them through the same saved preprocessing and XGBoost pipeline used for evaluation.

One persisted pipeline owns category encoding and prediction. The interface turns that output into a relative score, risk band, and review signal with explicit educational framing.

03 / How it works

How it works

The complete path from input to a finished, inspectable result.

Credit Risk Explorer workflow
Credit Risk Explorer pipeline showing eight inputs, validation, preprocessing, the XGBoost model, and an explained educational risk assessment.
  1. Collect 8 inputs

    Capture the applicant and loan fields used during training.

  2. Validate values

    Reject missing, invalid, or out-of-range input before inference.

  3. Apply preprocessing

    Use the saved column mapping and one-hot encoding.

  4. Run XGBoost

    Load the evaluated pipeline and calculate relative risk.

  5. Explain the output

    Show the score, risk band, review flag, and educational notice.

04 / What I built

What I built

The concrete parts I designed, implemented, and tested.

  1. Validated Streamlit interface

    I built the form and result workspace for the eight model inputs, including clear invalid-input and missing-model states.

  2. Saved inference pipeline

    I implemented one scikit-learn pipeline for preprocessing and XGBoost so training and the live interface use the same transformations.

  3. Repeatable training and evaluation

    I created the deterministic stratified split, model training, saved artifact, and evaluation output for accuracy, ROC-AUC, and bad-credit recall.

  4. Behavior-focused tests

    I added tests for field mapping, split isolation, saved-model round trips, validation, missing artifacts, and frontend/backend agreement.

05 / Results

Results

What the finished system demonstrates through working behavior, tests, and project artifacts.

  • A deterministic stratified 80/20 split keeps evaluation repeatable.
  • The held-out run records 0.756 ROC-AUC and 71.7% bad-credit recall on 200 records.
  • Nine tests cover data mapping, split isolation, saved-model round trips, validation, missing artifacts, and frontend/backend agreement.
  • The interface labels the result as a relative educational estimate rather than an approval decision.
Credit Risk Explorer assessment showing a 72 out of 100 relative risk score, a high risk band, and an additional-review signal.
The working assessment keeps the score, risk band, review signal, and educational explanation together in one inspectable result.