About CardioRisk AI
A comprehensive overview of the system architecture, clinical input biomarkers, Machine Learning algorithms, data preprocessing, model training, and prediction pipeline.
A. About The Project
CardioRisk AI is an applied Machine Learning healthcare web application designed to evaluate a patient's risk of developing Cardiovascular Disease (CVD). The system provides an accessible, non-invasive screening interface that estimates cardiac disease risk in real time based on routinely gathered health vitals and lifestyle indicators.
How Machine Learning is Used: The system feeds 11 standardized clinical vitals into a trained binary classification model. Rather than relying on rigid manual threshold rules, the model learns complex multivariate patterns and non-linear relationships from thousands of real-world patient records, estimating both the categorical risk class (Low vs. High Risk) and a calibrated probability score.
B. What is Cardiovascular Disease?
Cardiovascular Disease (CVD) refers to a broad spectrum of disorders affecting the heart and blood vessels. These include coronary artery disease, cerebrovascular disease, hypertension, peripheral arterial disease, and heart failure.
Why it is Important: Cardiovascular diseases are the leading cause of mortality globally. Many cardiac conditions develop silently over several years through the gradual accumulation of fatty deposits (atherosclerosis) inside arterial walls without noticeable early symptoms.
Why Early Risk Prediction is Useful: When identified at an early or pre-symptomatic stage, cardiovascular risk can often be substantially reduced or reversed through proactive lifestyle modifications, dietary changes, regular aerobic activity, and medical therapies. Automated risk prediction empowers individuals and practitioners to intervene before acute cardiovascular events occur.
Primary Contributing Factors: Key risk determinants include advancing age, elevated systolic and diastolic blood pressure, elevated blood cholesterol and glucose levels, excessive body mass index (BMI), tobacco smoking, alcohol abuse, and physical inactivity.
C. Actual Input Parameters
Our prediction system evaluates exactly 11 clinical parameters corresponding to the features utilized during model training. No parameters are assumed or fabricated; each represents an established cardiovascular indicator:
Age
Patient chronological age in years. Vascular elasticity decreases and arterial plaque accumulates with age, making it one of the strongest non-modifiable cardiovascular risk factors.
Gender
Biological sex. Men statistically exhibit earlier onset of coronary artery disease, whereas cardiovascular risk in females rises notably post-menopause due to hormonal shifts.
Height
Standing height in centimeters. Combined with weight to calculate Body Mass Index (BMI) and evaluate proportional body surface area.
Weight
Total body mass in kilograms. Essential for computing BMI and identifying obesity, which strains myocardial workload and correlates with metabolic syndrome.
Systolic Blood Pressure
Peak pressure exerted against arterial walls during ventricular contraction. Values above 120 mmHg indicate pre-hypertension, while sustained values ≥ 130 mmHg directly damage arterial lining.
Diastolic Blood Pressure
Minimum resting pressure in the arterial system between heartbeats. Elevated diastolic pressure (≥ 80 mmHg) indicates chronic peripheral vascular resistance.
Cholesterol Level
Blood cholesterol category. Hypercholesterolemia leads to lipid deposition in vessel lumens, accelerating atheroma formation and arterial stenosis.
Glucose Level
Fasting blood sugar category. Elevated blood glucose promotes endothelial inflammation, accelerates vessel stiffening, and is a hallmark of diabetic cardiomyopathy.
Smoking Status
Whether the patient regularly smokes tobacco. Inhaled toxins damage arterial endothelium, increase platelet aggregation, and induce acute coronary vasoconstriction.
Alcohol Intake
Regular consumption of alcoholic beverages. Chronic intake contributes to elevated systemic blood pressure, secondary cardiomyopathy, and cardiac arrhythmias.
Physical Activity
Engagement in regular physical exercise. Active individuals exhibit superior myocardial efficiency, improved lipid profiles, and reduced cardiovascular mortality.
D. Machine Learning Algorithm Used
Based on our project's codebase (backend/train_model.py and serialized model cardio_model.pkl), the machine learning model deployed in this project is:
What the Algorithm Is: Logistic Regression is a foundational supervised Machine Learning classification algorithm designed to model the probability of a binary outcome (here, the presence or absence of cardiovascular disease: cardio ∈ {0, 1}).
P(Cardio = 1 | x) = σ(z) = 1 / (1 + e-z)
- Why it is used in this project: Logistic Regression provides outstanding interpretability, computational efficiency, and direct probability calibration. Unlike black-box models, it yields an explicit mathematical probability score that directly translates to patient risk assessment.
- What it does with input data: The model computes a weighted linear sum of the 11 standardized clinical features. This linear score \(z\) is passed through the Sigmoid activation function \(\sigma(z)\), mapping the output strictly to a continuous probability interval between 0.0 and 1.0 (0% to 100%).
- How it helps in prediction: By applying a standard clinical decision boundary at threshold 0.50 (50%), any patient with \(P \ge 0.50\) is classified as High Risk, while patients with \(P < 0.50\) are classified as Low Risk.
-
Hyperparameters used in code: Trained with
max_iter=2000to ensure complete numerical convergence, andrandom_state=42for full scientific reproducibility. - Actual Evaluation Performance: Achieves a consistent test accuracy of ~72.6% to 73.0% on unseen patient records, demonstrating balanced precision and recall across both healthy and at-risk cohorts.
E. Data Preprocessing and Scaling
Effective machine learning requires rigorous preprocessing of raw clinical records before model fitting. The actual techniques applied in our codebase are:
Standard Feature Scaling (StandardScaler)
Scaling Method: StandardScaler from sklearn.preprocessing, serialized as scaler.pkl.
- What it does: Standardizes features by subtracting the empirical mean \(\mu\) and dividing by the standard deviation \(\sigma\), centering each feature at mean 0 with unit variance 1.
- Why scaling is required: The 11 input parameters have radically different measurement scales (e.g. age: 30–65 years, systolic BP: 100–200 mmHg, height: 150–190 cm, compared to binary flags like smoke or alcohol: 0 or 1). Without scaling, features with large numerical magnitudes would artificially dominate gradient updates and bias the model weights.
- Effect on the model: Ensures uniform gradient descent convergence, prevents coefficient distortion, and guarantees that every biomarker contributes proportionately based on its true diagnostic relevance.
Data Cleaning & Unit Transformation
-
Age Conversion: In the raw dataset, age was originally recorded in days. The training script converts days into standard medical years using
df['age'] = np.ceil(df['age'] / 365).astype(int), aligning it with the years entered on the frontend. - Blood Pressure Outlier Filtering: Extreme erroneous physiological values caused by sensor or measurement artifacts were filtered: systolic pressure constrained to \(60 \le ap\_hi \le 250\) mmHg and diastolic pressure to \(40 \le ap\_lo \le 200\) mmHg.
-
Feature Column Selection: The arbitrary patient identifier column
idwas removed from the feature matrix so that the model learns purely from diagnostic biomarkers.
F. Model Training Process
The machine learning training pipeline implemented in backend/train_model.py follows standard clinical data science methodology:
-
Dataset: Trained on the widely recognized Cardiovascular Disease Dataset (
cardio_train.csv), comprising 70,000 comprehensive patient records. -
Train / Test Split: The cleaned dataset was partitioned into an 80% training set (56,000 samples) and a 20% testing set (14,000 samples) using
test_size=0.20withstratify=Y(preserving equal proportion of disease-positive and disease-negative samples across splits) andrandom_state=42. -
Fitting the Model: The
StandardScalerwas fitted exclusively onX_trainto prevent data leakage, and bothX_trainandX_testwere scaled. The Logistic Regression estimator learned the optimal coefficient weights \(\beta\) using maximum likelihood estimation. -
Evaluation: Evaluated on the held-out 20% test dataset using Scikit-Learn's
accuracy_scoreand comprehensiveclassification_report(assessing Precision, Recall, and F1-Score). -
Serialization: The fitted estimator and scaler were saved using
joblib.dumpascardio_model.pklandscaler.pklfor zero-latency inference during web API requests.
G. Step-by-Step Prediction Process
When a patient or practitioner performs an assessment, data traverses the following end-to-end Machine Learning inference pipeline:
Step 1: Patient Enters Health Metrics
The user inputs age, gender, height, weight, systolic/diastolic blood pressure, cholesterol, glucose, and lifestyle factors on the Check Up page. Live BMI preview updates automatically.
Step 2: Transmission to Flask REST API
Frontend JavaScript validates the fields and sends an asynchronous JSON POST payload to the backend endpoint /api/predict.
Step 3: StandardScaler Feature Transformation
The backend validates data bounds and transforms the 11-element feature vector using the pre-fitted scaler.pkl, centering and scaling each biomarker.
Step 4: Logistic Regression Model Inference
The scaled vector is evaluated by cardio_model.pkl. The model computes log-odds and passes them through the Sigmoid function to derive class probability.
Step 5: Clinical Metric Computation
The backend calculates BMI and category, evaluates Blood Pressure stage (AHA classification), and generates targeted lifestyle and medical recommendations.
Step 6: Comprehensive Result Rendered to User
The frontend renders an animated risk gauge, color-coded risk badge (Low Risk in emerald green vs. High Risk in rose red), BMI & BP status cards, and personalized recommendations.