Project Documentation

About CardioRisk AI

A comprehensive overview of the system architecture, clinical input biomarkers, Machine Learning algorithms, data preprocessing, model training, and prediction pipeline.

A. About The Project

CardioRisk AI is an applied Machine Learning healthcare web application designed to evaluate a patient's risk of developing Cardiovascular Disease (CVD). The system provides an accessible, non-invasive screening interface that estimates cardiac disease risk in real time based on routinely gathered health vitals and lifestyle indicators.

Project Purpose: Traditional cardiac diagnostic screenings can be costly, time-consuming, and inaccessible for routine check-ups. This project demonstrates how modern supervised Machine Learning algorithms can assist clinical decision-making by rapidly identifying high-risk individuals from standard biometric data, encouraging timely consultation with certified healthcare professionals.

How Machine Learning is Used: The system feeds 11 standardized clinical vitals into a trained binary classification model. Rather than relying on rigid manual threshold rules, the model learns complex multivariate patterns and non-linear relationships from thousands of real-world patient records, estimating both the categorical risk class (Low vs. High Risk) and a calibrated probability score.

B. What is Cardiovascular Disease?

Cardiovascular Disease (CVD) refers to a broad spectrum of disorders affecting the heart and blood vessels. These include coronary artery disease, cerebrovascular disease, hypertension, peripheral arterial disease, and heart failure.

Why it is Important: Cardiovascular diseases are the leading cause of mortality globally. Many cardiac conditions develop silently over several years through the gradual accumulation of fatty deposits (atherosclerosis) inside arterial walls without noticeable early symptoms.

Why Early Risk Prediction is Useful: When identified at an early or pre-symptomatic stage, cardiovascular risk can often be substantially reduced or reversed through proactive lifestyle modifications, dietary changes, regular aerobic activity, and medical therapies. Automated risk prediction empowers individuals and practitioners to intervene before acute cardiovascular events occur.

Primary Contributing Factors: Key risk determinants include advancing age, elevated systolic and diastolic blood pressure, elevated blood cholesterol and glucose levels, excessive body mass index (BMI), tobacco smoking, alcohol abuse, and physical inactivity.

C. Actual Input Parameters

Our prediction system evaluates exactly 11 clinical parameters corresponding to the features utilized during model training. No parameters are assumed or fabricated; each represents an established cardiovascular indicator:

age
Demographic • Integer (Years)

Age

Patient chronological age in years. Vascular elasticity decreases and arterial plaque accumulates with age, making it one of the strongest non-modifiable cardiovascular risk factors.

gender
Demographic • Binary (1: Female, 2: Male)

Gender

Biological sex. Men statistically exhibit earlier onset of coronary artery disease, whereas cardiovascular risk in females rises notably post-menopause due to hormonal shifts.

height
Biometric • Float (cm)

Height

Standing height in centimeters. Combined with weight to calculate Body Mass Index (BMI) and evaluate proportional body surface area.

weight
Biometric • Float (kg)

Weight

Total body mass in kilograms. Essential for computing BMI and identifying obesity, which strains myocardial workload and correlates with metabolic syndrome.

ap_hi
Clinical Vital • Integer (mmHg)

Systolic Blood Pressure

Peak pressure exerted against arterial walls during ventricular contraction. Values above 120 mmHg indicate pre-hypertension, while sustained values ≥ 130 mmHg directly damage arterial lining.

ap_lo
Clinical Vital • Integer (mmHg)

Diastolic Blood Pressure

Minimum resting pressure in the arterial system between heartbeats. Elevated diastolic pressure (≥ 80 mmHg) indicates chronic peripheral vascular resistance.

cholesterol
Biochemical • Ordinal (1: Normal, 2: Above, 3: Well Above)

Cholesterol Level

Blood cholesterol category. Hypercholesterolemia leads to lipid deposition in vessel lumens, accelerating atheroma formation and arterial stenosis.

gluc
Biochemical • Ordinal (1: Normal, 2: Above, 3: Well Above)

Glucose Level

Fasting blood sugar category. Elevated blood glucose promotes endothelial inflammation, accelerates vessel stiffening, and is a hallmark of diabetic cardiomyopathy.

smoke
Lifestyle • Binary (0: No, 1: Yes)

Smoking Status

Whether the patient regularly smokes tobacco. Inhaled toxins damage arterial endothelium, increase platelet aggregation, and induce acute coronary vasoconstriction.

alco
Lifestyle • Binary (0: No, 1: Yes)

Alcohol Intake

Regular consumption of alcoholic beverages. Chronic intake contributes to elevated systemic blood pressure, secondary cardiomyopathy, and cardiac arrhythmias.

active
Lifestyle • Binary (0: Inactive, 1: Active)

Physical Activity

Engagement in regular physical exercise. Active individuals exhibit superior myocardial efficiency, improved lipid profiles, and reduced cardiovascular mortality.

D. Machine Learning Algorithm Used

Based on our project's codebase (backend/train_model.py and serialized model cardio_model.pkl), the machine learning model deployed in this project is:

Logistic Regression (Scikit-Learn)
Test Accuracy: ~72.6% – 73.0%

What the Algorithm Is: Logistic Regression is a foundational supervised Machine Learning classification algorithm designed to model the probability of a binary outcome (here, the presence or absence of cardiovascular disease: cardio ∈ {0, 1}).

z = β0 + β1x1 + β2x2 + … + β11x11
P(Cardio = 1 | x) = σ(z) = 1 / (1 + e-z)
  • Why it is used in this project: Logistic Regression provides outstanding interpretability, computational efficiency, and direct probability calibration. Unlike black-box models, it yields an explicit mathematical probability score that directly translates to patient risk assessment.
  • What it does with input data: The model computes a weighted linear sum of the 11 standardized clinical features. This linear score \(z\) is passed through the Sigmoid activation function \(\sigma(z)\), mapping the output strictly to a continuous probability interval between 0.0 and 1.0 (0% to 100%).
  • How it helps in prediction: By applying a standard clinical decision boundary at threshold 0.50 (50%), any patient with \(P \ge 0.50\) is classified as High Risk, while patients with \(P < 0.50\) are classified as Low Risk.
  • Hyperparameters used in code: Trained with max_iter=2000 to ensure complete numerical convergence, and random_state=42 for full scientific reproducibility.
  • Actual Evaluation Performance: Achieves a consistent test accuracy of ~72.6% to 73.0% on unseen patient records, demonstrating balanced precision and recall across both healthy and at-risk cohorts.

E. Data Preprocessing and Scaling

Effective machine learning requires rigorous preprocessing of raw clinical records before model fitting. The actual techniques applied in our codebase are:

Standard Feature Scaling (StandardScaler)

Scaling Method: StandardScaler from sklearn.preprocessing, serialized as scaler.pkl.

z = (x - μ) / σ
  • What it does: Standardizes features by subtracting the empirical mean \(\mu\) and dividing by the standard deviation \(\sigma\), centering each feature at mean 0 with unit variance 1.
  • Why scaling is required: The 11 input parameters have radically different measurement scales (e.g. age: 30–65 years, systolic BP: 100–200 mmHg, height: 150–190 cm, compared to binary flags like smoke or alcohol: 0 or 1). Without scaling, features with large numerical magnitudes would artificially dominate gradient updates and bias the model weights.
  • Effect on the model: Ensures uniform gradient descent convergence, prevents coefficient distortion, and guarantees that every biomarker contributes proportionately based on its true diagnostic relevance.

Data Cleaning & Unit Transformation

  • Age Conversion: In the raw dataset, age was originally recorded in days. The training script converts days into standard medical years using df['age'] = np.ceil(df['age'] / 365).astype(int), aligning it with the years entered on the frontend.
  • Blood Pressure Outlier Filtering: Extreme erroneous physiological values caused by sensor or measurement artifacts were filtered: systolic pressure constrained to \(60 \le ap\_hi \le 250\) mmHg and diastolic pressure to \(40 \le ap\_lo \le 200\) mmHg.
  • Feature Column Selection: The arbitrary patient identifier column id was removed from the feature matrix so that the model learns purely from diagnostic biomarkers.

F. Model Training Process

The machine learning training pipeline implemented in backend/train_model.py follows standard clinical data science methodology:

  • Dataset: Trained on the widely recognized Cardiovascular Disease Dataset (cardio_train.csv), comprising 70,000 comprehensive patient records.
  • Train / Test Split: The cleaned dataset was partitioned into an 80% training set (56,000 samples) and a 20% testing set (14,000 samples) using test_size=0.20 with stratify=Y (preserving equal proportion of disease-positive and disease-negative samples across splits) and random_state=42.
  • Fitting the Model: The StandardScaler was fitted exclusively on X_train to prevent data leakage, and both X_train and X_test were scaled. The Logistic Regression estimator learned the optimal coefficient weights \(\beta\) using maximum likelihood estimation.
  • Evaluation: Evaluated on the held-out 20% test dataset using Scikit-Learn's accuracy_score and comprehensive classification_report (assessing Precision, Recall, and F1-Score).
  • Serialization: The fitted estimator and scaler were saved using joblib.dump as cardio_model.pkl and scaler.pkl for zero-latency inference during web API requests.

G. Step-by-Step Prediction Process

When a patient or practitioner performs an assessment, data traverses the following end-to-end Machine Learning inference pipeline:

Step 1: Patient Enters Health Metrics

The user inputs age, gender, height, weight, systolic/diastolic blood pressure, cholesterol, glucose, and lifestyle factors on the Check Up page. Live BMI preview updates automatically.

Step 2: Transmission to Flask REST API

Frontend JavaScript validates the fields and sends an asynchronous JSON POST payload to the backend endpoint /api/predict.

Step 3: StandardScaler Feature Transformation

The backend validates data bounds and transforms the 11-element feature vector using the pre-fitted scaler.pkl, centering and scaling each biomarker.

Step 4: Logistic Regression Model Inference

The scaled vector is evaluated by cardio_model.pkl. The model computes log-odds and passes them through the Sigmoid function to derive class probability.

Step 5: Clinical Metric Computation

The backend calculates BMI and category, evaluates Blood Pressure stage (AHA classification), and generates targeted lifestyle and medical recommendations.

Step 6: Comprehensive Result Rendered to User

The frontend renders an animated risk gauge, color-coded risk badge (Low Risk in emerald green vs. High Risk in rose red), BMI & BP status cards, and personalized recommendations.

Try the Prediction System Now