Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 9s (2025) 589 https://internationalpubls.com Ensemble Learning based Cardiovascular Disease Prediction Combining with Predictive Analytics & Risk Stratification S. Ramchandra Reddy 1,*, G. Vishnu Murthy 2 1 Research scholar, Department of Computer Science and Engineering, Anurag University, Hyderabad, Tealangana, India, 500088. 2 Professor, Department of Computer Science and Engineering, Anurag University, Hyderabad, Tealangana, India, 500088. Author Emails: rcreddy79@gmail.com1, deancse@anurag.edu.in2 Article History: Received: 15-11-2024 Revised: 26-12-2024 Accepted: 10-01-2025 Abstract: Heart failure continues to be a significant contributor to illness and death on a global scale, highlighting the need for effective strategies in early detection and risk stratification. Therefore, the study presents an efficient approach to heart failure prediction by integrating predictive analytics with advanced risk stratification techniques. This study develops and evaluates a novel ensemble learning-based predictive model that combines machine learning algorithms with clinical and demographic data to enhance early diagnosis and risk assessment. The proposed model leverages multiple data sources, such as electronic health records and patient history, to identify critical predictors and generate actionable insights. This study demonstrates significant improvements in prediction accuracy compared to GBC, RF, LR, DTC, and MLP 4%, 4%, 6%, 5%, and 4%, respectively, with enhanced capability to identify high-risk individuals before the onset of severe symptoms. By integrating these predictive and stratification techniques, the study offers a robust framework for early intervention and personalized treatment, ultimately contributing to better management of heart failure and improved patient outcomes. These findings show the potential of combining advanced analytics with clinical expertise to advance heart failure prediction and management. Keywords: Cardiovascular Disease, Electronic Health Record(EHR), Ensemble, Heart Failure, Machine Learning, Random Forest. 1. Introduction Heart failure (HF) is a complex and critical public health issue characterized by the heart's inability to pump blood efficiently, leading to significant morbidity, mortality, and a heavy burden on healthcare systems worldwide [1]. HF affects millions of individuals, and its prevalence continues to rise due to factors such as aging populations, lifestyle changes, and increasing rates of cardiovascular diseases. This growing global health concern underscores the need for innovative approaches to early detection, accurate diagnosis, and effective management of HF. Traditional diagnostic methods, which often identify HF at its later stages when symptoms are more severe, limit the potential for early treatment and intervention, resulting in poorer patient outcomes and higher healthcare costs [2], [3]. Recent advancements in predictive analytics and machine learning present promising opportunities to address these challenges by enhancing the accuracy and timeliness of HF prediction [4], [5]. Predictive models that leverage diverse data sources, including electronic health records (EHRs), patient mailto:rcreddy79@gmail.com1 mailto:deancse@anurag.edu.in Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 9s (2025) 590 https://internationalpubls.com demographics, clinical history, and lifestyle data, offer a more comprehensive assessment of a patient’s risk profile. These models are capable of identifying patterns, correlations, and risk factors that traditional diagnostic methods may overlook. By employing machine learning algorithms, it becomes possible to predict HF at an earlier stage, thus enabling clinicians to implement preventive measures and more personalized treatment plans. Early detection through predictive analytics can potentially reduce the progression of HF and improve the quality of life for patients. A key component of effective HF management is risk stratification, which involves categorizing patients based on their likelihood of developing HF or experiencing adverse outcomes. Accurate risk stratification is essential for healthcare providers to prioritize clinical interventions, allocate resources efficiently, and tailor treatment plans to individual patient needs [6]. Current risk stratification techniques often lack the precision required for personalized care and may not fully utilize the vast amount of available healthcare data. As a result, there is a need for more sophisticated models that integrate advanced analytics with robust risk stratification methods. The objective of this research is to develop and evaluate an integrated approach that combines predictive analytics with ensemble learning techniques to enhance HF prediction and risk stratification. Ensemble learning, which aggregates the predictive capabilities of multiple machine learning models, has been shown to improve accuracy and robustness in various predictive tasks. By leveraging ensemble methods, this study aims to create a predictive model that not only outperforms traditional diagnostic approaches but also delivers more reliable and actionable insights into patient risk. The proposed model will utilize a broad array of clinical and demographic data to build a holistic framework for HF prediction, enabling earlier identification of high-risk individuals and supporting personalized treatment strategies. This research also seeks to bridge the gap between advancements in machine learning and their practical application in real-world clinical settings. Despite the potential of predictive analytics, its integration into routine clinical practice for HF prediction remains limited. Many predictive models are developed in research settings but are not widely adopted due to challenges related to implementation, validation, and clinician usability. Therefore, this study focuses not only on the development of a predictive model but also on ensuring its practical applicability and scalability in healthcare environments. By integrating predictive analytics with advanced risk stratification techniques, this study aims to contribute to more effective and personalized HF management. The proposed model will support earlier intervention, allowing healthcare providers to implement preventive measures before the onset of severe symptoms, ultimately improving patient outcomes and optimizing resource allocation within healthcare systems. This research has the potential to significantly advance the field of HF prediction by providing a comprehensive, data-driven framework that enhances predictive accuracy, facilitates early diagnosis, and enables personalized care. 2. Method Research in HF prediction has evolved significantly, leveraging various methodologies and technologies to enhance early detection and management. This section reviews key studies and methodologies related to predictive analytics and risk stratification in HF. Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 9s (2025) 591 https://internationalpubls.com 2.1. Traditional Predictive Models Early HF prediction methods primarily relied on statistical models and risk scoring systems [7], [8]. Widely used models, such as the Framingham Risk Score [9] and the Chicago Heart Association (CHA) risk score [10], estimate HF likelihood based on traditional risk factors like age, hypertension, and diabetes. However, these models often lack the precision needed for early detection and personalized risk assessment. 2.2. Machine Learning (ML) in HF Prediction Recent advancements in ML have improved HF prediction accuracy. Various models, such as decision trees (DTs), support vector machines (SVMs), and XGBoost classifiers (XGBC), analyse large datasets to identify complex patterns associated with HF [11], [12]. For example, Ganie et al. [13] showed that ensemble methods and gradient boosting significantly enhance prediction performance over traditional models. 2.3. Deep Learning (DL) in HF Prediction DL models, including convolutional neural networks (CNNs) and recurrent neural networks (RNNs), have advanced HF prediction further [14], [15], [16]. Guo et al. [17] demonstrated that deep neural networks analysing electronic health records (EHRs) improve HF onset prediction accuracy by capturing complex relationships within extensive datasets. 2.4. Risk Stratification and Personalization Risk stratification models aim to categorize patients based on their likelihood of risk. These models categorize patients based on HF risk to enable targeted interventions. Sophisticated algorithms, such as clustering and Bayesian networks [18], have refined this process. Strianese et al. [19] showed that integrating clinical data with genetic information enhances risk stratification and personalized treatment. 2.5. Integration of Predictive Analytics and Clinical Practice Despite advancements, there is a gap in integrating predictive models into clinical Despite advancements, integrating predictive models into clinical practice remains challenging. Azizi et al. [20] noted difficulties in translating predictive analytics into actionable clinical tools, highlighting the need for models that are both accurate and practical for real-world use. 2.6. Comprehensive Approaches and Novel Frameworks Recent studies advocate for frameworks that combine predictive analytics with risk stratification for comprehensive HF management [21], [22]. Adewole et al. [23] proposed a model that integrates predictive analytics with real-time monitoring to enhance early detection and personalized risk assessment. This research advances these methods by introducing a novel approach that merges advanced machine learning with risk stratification to boost HF prediction accuracy and clinical relevance. 3. Materials and Methods 3.1. Model architecture The proposed model architecture integrates predictive analytics with risk stratification to enhance HF prediction described in Figure 1. This architecture consists of the following key components: Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 9s (2025) 592 https://internationalpubls.com 3.1.1. Data Input Layer The model begins with an input layer that accepts diverse data sources, including EHRs, clinical metrics, and patient demographics. Let X represent the input feature matrix with n samples, where each sample xi Rm includes features such as age, blood pressure, and lab results. Figure 1. Model Architecture of the Proposed Model 3.1.2. Feature Embedding To handle the heterogeneous nature of the input features, an embedding layer transforms categorical features into continuous vectors. This process captures the underlying relationships and interactions between different features. For a categorical feature xi c, the embedding function  maps it to a continuous vector. ei c =  (xi c) (1) where, ei c   Rd is the embedding vector. 3.1.3. Feature Integration The embedded features are concatenated with continuous features to form a unified feature vector hi hi = concat ( xi c, xi c, ei c) (2) 3.1.4. Ensemble Network The integrated feature vector hi is passed through a series of layers within the ensemble network, enabling the proposed model to learn complex patterns and interactions. The architecture consists of: hi → Layer1 → Layer2 →….→ Layern (3) Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 9s (2025) 593 https://internationalpubls.com 3.1.5. Output Layer The final layer is a logistic regression layer that outputs the probability P(yi = 1| xi ) of HF for each patient. The output is computed using a sigmoid activation function: �̂�𝒾 = 𝜎(𝑤𝑜𝒉𝒾 (2) + 𝑏𝑜) (4) where,  is the sigmoid function, wo is the weight matrix, and bo is the bias term. 3.1.6. Risk Stratification Based on the predicted probabilities �̂�𝑖, patients are classified into different risk categories. The risk stratification is performed using thresholds τk defining categories Ck as: 𝐶𝑘 = {𝒙𝒾 | 𝜏𝑘−1 < �̂�𝒾 ≤ 𝜏𝑘 } (5) 3.2. Problem Formulation HF prediction involves estimating the probability that a patient will develop HF based on a range of clinical and demographic variables. The problem can be formulated as a classification task where the goal is to predict  binary outcome, y, indicating the presence or absence of HF. Let X represent the feature matrix consisting of n patients, where each patient i is described by a vector of features xi Rm . The goal is to model the probability P(yi = 1| xi ) where yi is a binary indicator of HF for patient i. To achieve this, we define a predictive model f(xi;θ) that outputs the probability of HF given the features xi and parameters θ : 𝑃(𝑦𝒾 = 1|𝑥𝑖) = 𝑓(𝑥𝒾; 𝜃) (6) where f(xi;θ) is typically a logistic regression function, neural network, or any other classification model. The model is trained to minimize the loss function, which is often the binary cross-entropy loss for classification problems: 𝐿(𝜃) = − 1 𝑛 ∑ [𝑦𝒾𝑙𝑜𝑔(𝑓(𝑥𝒾; 𝜃)) + (1 − 𝑦𝒾) 𝑙𝑜𝑔 (1 − 𝑓(𝑥𝒾; 𝜃))]𝑛 𝒾=1 (7) In addition to prediction, risk stratification is employed to classify patients into different risk categories based on their predicted probabilities. Let �̂�𝒾 denote the predicted probability for patient i. Patients are classified into risk categories Ck (e.g., low, medium, high risk) based on thresholds τk: Ck ={xi |τk−1