Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 9s (2025) 2754 https://internationalpubls.com Probabilistic Perspective of Jungck Contractive Situations in Machine Learning Ashutosh Mishra1, Piyush Kumar Tripathi2, Mukesh Kumar Singh3, Manisha Gupta4, Alok Kumar Agrawal5 1,2,5Department of Mathematics, ASAS, Amity University Uttar Pradesh (Lucknow Campus), India 3Armapore P. G. College, Kanpur, India 4University of Technology and Applied Science, Muscat Email: 1amishra94154@gmail.com, 1ashutosh.mishra8@s.amity.edu; 2pktripathi@amity.edu; 3mksinghuptti@gmail.com; 4manisha.gupta@utas.edu.om; 5akagrawal@lko.amity.edu Article History: Received: 12-01-2025 Revised: 15-02-2025 Accepted: 01-03-2025 Abstract: In the realm of machine learning (ML), the underlying mathematical frameworks significantly influence the development and efficacy of algorithms. One such framework is the probabilistic metric space, which provides a robust foundation for handling uncertainty and variability inherent in real-world problems and data. This work explores the applicability of probabilistic metric spaces in machine learning, with a particular focus on the Jungck contractive notion. The Jungck contraction theorem extends the classical Banach contraction principle to the setting of probabilistic metric spaces. This theorem is pivotal for proving the existence and uniqueness of fixed points in probabilistic settings, which is crucial for the convergence analysis of ML algorithms. This generalises invariant point propositions to probabilistic settings, offers powerful tools for the convergence analysis of iterative ML algorithms. We delve into the theoretical underpinnings of probabilistic metric spaces, explain the Jungck contractive notion, and bespeak its relevance and application in ML algorithms. Keywords: Probabilistic space, Contractive conditions, Machine learning, Convergence analysis. 1. Introduction Machine learning is an underlying collection as a subset of AI (artificial intelligence) that concentrates the advancement of algorithms, as well as statistical model which enables computers to execute certain tasks without distinct instructions, depending on patterns and inference instead. In other words, it is about creating systems that can learn from data. The crux behind such learning is to empower computer related programmes to train automatically from given data and to improve its effectiveness over time, without being distinctively programmed for every task. This is brought about through various learning techniques viz. supervised, unsupervised, reinforced, and deep. Supervised Learning In this approach, the algorithm retains information from some specific labelled data, meaning there by that it is given input-output pairs and learns to map inputs to outputs. For example, let us have a given set of labelled impressions (images) of cats and dogs, a supervised learning algorithm would learn to classify new impressions (images) as either cats or dogs [8]. mailto:amishra94154@gmail.com mailto:1ashutosh.mishra8@s.amity.edu mailto:2pktripathi@amity.edu mailto:mksinghuptti@gmail.com3 mailto:manisha.gupta@utas.edu.om Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 9s (2025) 2755 https://internationalpubls.com Unsupervised Learning In this stage, the algorithm is given input data without explicit labels, and it tries to find patterns or structure in the data. Clustering is a common task in unsupervised learning, where the algorithm groups similar data points together. An example could be grouping similar customer purchase behaviour for market segmentation [8]. Reinforced Learning This is decision making training algorithms that gets learning from feedback. The algorithm trains itself to achieve a outcome or goal in an uncertain, potentially complex environment by taking actions and receiving rewards or penalties. It is often used in areas like gamming, robotics, and auto driving [10]. Deep Learning It is an underlying subset of machine learning which uses artificial neural networks with many blankets of layers, hence proposed as β€œdeep”, to model and process complex patterns in large amounts of data. It's particularly effective for tasks such as image and speech recognition, natural language processing, and playing strategic games like Go [9]. Overall, machine learning has numerous applications across various industries, including finance, e- commerce, healthcare, cyber security, and many more. Its ability to find patterns in data and make predictions or decisions based on those patterns has made it a powerful tool for solving a wide range of real-world problems. Machine learning algorithms often rely on the convergence properties of iterative processes. Classical metric spaces and their associated invariant point theorems have provided substantial insights into these processes. However, real-world data is frequently stochastic in nature, demanding a probabilistic approach to better model the uncertainty and variability. Probabilistic metric spaces, introduced by Menger[2]in the year 1940, extend the concept of metric spaces by incorporating probability distributions instead of deterministic distances. Since then, a variety of propositions were introduced using Menger theory. Some implications of this notion are illustrated in population dynamics of cells [11, 14]. This probabilistic framework is particularly useful in ML, where data and models are inherently noisy and probabilistic. The Jungck theorem on contraction mapping[4], a generalization of Banach's fixed point result to probabilistic metric spaces, offers a robust mechanism to analyze the convergence of sequences in these spaces[1]. This theorem is instrumental in ensuring the stability and reliability of iterative ML algorithms. By leveraging the Jungck contraction theorem, we can develop and analyze ML algorithms that are more resilient to the uncertainties of real-world data. The K-means clustering algorithm is a fundamental technique in machine learning for partitioning data into clusters based on similarity[7]. In traditional K-means, the distance between points and centroids is deterministic. However, real-world data often contains uncertainty and noise, which can lead to instability in the clustering results. By incorporating a probabilistic approach into K-means, we can account for this uncertainty, leading to more robust and reliable clustering. Before delving into the probabilistic K-means algorithm, it's essential to understand probabilistic metric spaces (PMS). Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 9s (2025) 2756 https://internationalpubls.com 2. Objectives By employing the Jungck contractive statement, we ensure that the probabilistic gradient descent algorithm converges to an optimal solution, providing robustness against the uncertainties in the training data. 3. Methods and Preliminaries A. Probabilistic Space: A probabilistic distance (or metric) space (PMS) is a generalization of a metric space where the distance between any two points is not a single real number but a distribution function. Formally, a PMS is a triplet (𝑃, 𝐹, 𝑇), where 𝑃 is a set, 𝐹 is the set of distribution functions on [0, ∞), and 𝑇: 𝑃 Γ— 𝑃 β†’ 𝐹is a mapping that satisfies certain axioms analogous to those of a metric space. Definition 3.1: [11] A distribution function 𝐹 on [0, ∞) is a non-decreasing, left-continuous function, such that, 𝐹(0) = 0, and π‘™π‘–π‘š π‘›β†’βˆž 𝐹(π‘₯) = 1. The function 𝑇(π‘₯, 𝑦) gives the probability that the distance between π‘₯ and 𝑦is less than or equal to a given value. This probabilistic approach is particularly useful in modelling uncertainties in data points. Definition 3.2: [11] A triangular norm (t-norm) can be represented as a mapping 𝑑: [0,1] Γ— [0,1] β†’ [0,1] satisfying: i. t(π‘Ž, 1) = π‘Ž for every π‘Ž ∈ [0,1] ii. t(π‘Ž, 𝑏) = t(𝑏, π‘Ž) for every π‘Ž, 𝑏 ∈ [0,1] iii. t(π‘Ž, 𝑐) = t(𝑏, 𝑑) if π‘Ž β‰₯ 𝑏, 𝑐 β‰₯ 𝑑 iv. t(π‘Ž, t(𝑏, 𝑐)) = t(t(π‘Ž, 𝑏), 𝑐); π‘Ž, 𝑏,𝑐 ∈ [0,1]. B. Axioms of PMS: i. Non-negativity: 𝑇(π‘₯, 𝑦) = πœ–0 (where πœ–0 is the degenerate distribution at 0), iff, π‘₯ = 𝑦. ii. Symmetry: 𝑇(π‘₯, 𝑦) = 𝑇(𝑦, π‘₯), π‘“π‘œπ‘Ÿ π‘Žπ‘™π‘™ π‘₯ ∈ π‘ƒπ‘Žπ‘›π‘‘ 𝑦 ∈ 𝑃. iii. Triangle inequality: For all π‘₯, 𝑦, 𝑧 in 𝑃 and for all 𝑑 β‰₯ 0, 𝑇(π‘₯, 𝑧)(𝑑) β‰₯ sup𝑒+𝑣=𝑑 π‘šπ‘–π‘› {𝑇(π‘₯, 𝑦)(𝑒), 𝑇(𝑦, 𝑧)(𝑣)}. These axioms ensure that the probabilistic metric retains the essential properties of a classical metric, allowing us to extend various analytical techniques to the probabilistic domain. C. Jungck Contractive Condition: The Jungck contraction theorem extends the classical Banach contraction principle to the setting of probabilistic metric spaces. This theorem is pivotal for proving the existence and uniqueness of fixed points in probabilistic settings, which is crucial for the convergence analysis of ML algorithms. Theorem 3.1 (Jungck Contraction Theorem): [4] Let (𝑃, 𝐹, 𝑇) be a complete probabilistic metric space. Suppose 𝑓, 𝑔: 𝑃 β†’ 𝑃are two mappings satisfying: 𝑇(𝑓(π‘₯), 𝑓(𝑦))(𝑑) β‰₯ 𝛽 𝑇(𝑔(π‘₯), 𝑔(𝑦))(𝑑), for all π‘₯, 𝑦 ∈ 𝑃 and for some constant 𝛽 ∈ (0,1).Then,𝑓 and 𝑔 have a unique common invariant point in 𝑃. This result ensures that under a contraction condition, iterative applications of the mappings 𝑓 and 𝑔 will converge to a common invariant point. This result is useful in analyzing machine learning based algorithms where such mappings often represent iterative update steps, for example, in analysing the convergence of probabilistic K-means clustering. Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 9s (2025) 2757 https://internationalpubls.com Application in Machine Learning: Many algorithms, in machine learning, can be viewed as iterative processes that seek to minimize an objective function or optimize a model. The convergence of these processes is crucial for the reliability and stability of the algorithms. By framing these iterative processes within the context of probabilistic metric spaces, we can leverage the Jungck contractive condition to assure convergence even in the presence of uncertainty. K-means Clustering in Probabilistic Space-A Case Study: The K-means clustering algorithm is a widely used method for partitioning data into clusters. However, in the presence of uncertainty and noise, the classical K-means algorithm may struggle to find stable clusters. By adopting a probabilistic approach, we can enhance the robustness of K-means clustering. Probabilistic Distance Metric: Instead of using a deterministic distance, we define a probabilistic distance between points. This distance is modelled as a distribution function capturing the uncertainty in the data. Probabilistic Centroid Update: The update step for centroids in the classical K-means is replaced by a probabilistic update. Given a cluster, the new centroid is chosen such that it minimizes the expected probabilistic distance to the points in the cluster. Convergence Analysis: Using the Jungck contraction theorem, we can analyse the convergence of the probabilistic K-means algorithm. By ensuring that the probabilistic distance satisfies the contraction condition, we can guarantee that the iterative updates of centroids will converge to a stable configuration. K-means Clustering Algorithm in Probabilistic Space: The probabilistic K-means algorithm [7] extends the classical K-means by incorporating probabilistic distances, which are represented by distribution functions. This approach provides a more robust clustering method in the presence of uncertainty and noise. Algorithm of K-means in probabilistic space is as follows: 1. Initialization: Randomly initialize π‘˜ probabilistic centroids, where each centroid is represented by a distribution function. 2. Assignment Step: Here, we assign each data point to the cluster whose centroid minimizes the probabilistic distance. 3. Update Step: For each and every cluster, update the centroid to minimize the expected probabilistic distance to the points in the cluster. 4. Convergence Check: Repeat the assignment and update steps until the centroids converge. Step-by-Step Explanation: 1. Initialization: Randomly select π‘˜ points from the dataset as initial centroids. Represent each centroid as a distribution function capturing the uncertainty. 2. Assignment Step: Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 9s (2025) 2758 https://internationalpubls.com For each data point π‘₯𝑖, compute the probabilistic distance to each centroid 𝐢𝑗 using the distribution function 𝑇(π‘₯𝑖 , 𝐢𝑗). Assign π‘₯𝑖to the cluster with the centroid that has the smallest expected probabilistic distance. 3. Update Step: For each cluster, 𝐢𝑗, update the centroid to the point that minimizes the expected probabilistic distance to all points in the cluster. This involves finding a new distribution function that best represents the central tendency of the points in the cluster. 4. Convergence Check: We check, if the centroids have stabilized (i.e., there is little to no change in the distribution functions of the centroids). If not so, we repeat the assignment and update steps until the convergence is obtained, e.g. Consider a simple 2D dataset with three clusters. The classical K-means algorithm may struggle to identify the clusters accurately due to noise and uncertainty in the data points. We apply the probabilistic K-means algorithm to handle this uncertainty. Dataset: Points: [(1.0, 2.0), (1.5, 1.8), (5.0, 8.0), (8.0, 8.0), (1.0, 0.6), (9.0, 11.0)] Initialization: Randomly initialize π‘˜ = 2 probabilistic centroids. Assume initial centroids are (1.0, 2.0) and (5.0, 8.0), represented by distribution functions with small variances. Iteration 1: Assignment Step Compute probabilistic distances between each point and the centroids using a Gaussian distribution function. Assign points to the nearest centroid based on expected probabilistic distance. Iteration 1: Update Step Update centroids by recalculating the distribution functions that minimize the expected probabilistic distance within each cluster. Iteration 1: Results Cluster 1: [(1.0, 2.0), (1.5, 1.8), (1.0, 0.6)] Cluster 2: [(5.0, 8.0), (8.0, 8.0), (9.0, 11.0)] New Centroids (distribution functions): Cluster 1: Mean = (1.17, 1.47), Variance = small Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 9s (2025) 2759 https://internationalpubls.com Cluster 2: Mean = (7.33, 9.0), Variance = small Iteration 2: Assignment Step Recompute probabilistic distances using updated centroids. Reassign points to the nearest centroid. Iteration 2: Update Step Recalculate centroids for the new clusters. Iteration 2: Results Cluster 1: [(1.0, 2.0), (1.5, 1.8), (1.0, 0.6)] Cluster 2: [(5.0, 8.0), (8.0, 8.0), (9.0, 11.0)] New Centroids (distribution functions): Cluster 1: Mean = (1.17, 1.47), Variance = very small Cluster 2: Mean = (7.33, 9.0), Variance = very small Convergence Check: Centroids have stabilized (no significant change in mean and variance of distribution functions). Algorithm converges. Final Clusters: Cluster 1: [(1.0, 2.0), (1.5, 1.8), (1.0, 0.6)] Cluster 2: [(5.0, 8.0), (8.0, 8.0), (9.0, 11.0)] By applying the Jungck contraction theorem, we ensure that the probabilistic K-means algorithm converges to a stable clustering configuration, providing robustness against the inherent uncertainties in the data. Probabilistic Neural Networks (PNN)-A Case Study: Neural networks (NNs) are a cornerstone of modern ML. Training a NN involves iteratively updating weights to minimize a loss function. In the presence of noisy data, a probabilistic framework can enhance the training process[13].PNNs and related probabilistic models are being applied to various domains such as healthcare (medical diagnosis and prognosis) and finance (risk assessment and fraud detection). These applications benefit from the ability of PNNs to handle uncertainty and complex patterns in data[13]. 1. Probabilistic Weight Updates: The weight updates are modelled probabilistically, capturing the uncertainty in the gradient estimates due to noisy data. 2. Probabilistic Loss Function: The loss function is defined in terms of probabilistic distances, providing a more robust objective for optimization [16]. Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 9s (2025) 2760 https://internationalpubls.com 3. Convergence Analysis: Using the Jungck contraction theorem, we analyse the convergence of the probabilistic weight updates [16]. By ensuring that the probabilistic distance satisfies the contraction condition, we can assure convergence to an optimal set of weights. Probabilistic Gradient Descent Algorithm: 1. Initialization: Randomly initialize weights probabilistically. 2. Forward Pass: Compute the output of the NN using the current probabilistic weights. 3. Loss Computation: Compute the probabilistic loss based on the output and the true labels. 4. Backward Pass: Compute the probabilistic gradient of the loss with respect to the weights. 5. Weight Update: Update the weights probabilistically based on the computed gradient. 6. Convergence Check: We repeat steps 2-5, until the loss converges. By employing the Jungck contractive statement and commutativity [15], we ensure that the probabilistic gradient descent algorithm converges to an optimal solution, providing robustness against the uncertainties in the training data. 4. Main Results & Statements We use the following as a result in harnessing the situation as stated above. Lemma 4.1: [6] If (𝑃, 𝐹, 𝑇) be a Menger generalized PM-space with a Cauchy sequence < 𝛼𝑛 > in 𝑃along with some 𝛽 ∈ (0,1) such that 𝐹𝛼𝑛,𝛼𝑛+1 (𝛽π‘₯) β‰₯ πΉπ›Όπ‘›βˆ’1,𝛼𝑛 (π‘₯) for all positive values of π‘₯. Also, if for π‘Ÿ > 1 and 𝑖 = 𝑛 π‘‘π‘œ ∞, we have lim π‘›β†’βˆž 𝑇𝑖 (𝐹𝛼0,𝛼1 (π‘Ÿπ‘–)) = 1 Therefore, we say the sequence < 𝛼𝑛 > is 𝐹-Cauchy. To get the common coincidence point for mappings under generalized Menger space for producing some machine learning results, we can apply the following statements with some specific constraints. Theorem 4.1: With (𝑃, 𝐹, 𝑇)a generalized PM-space and a continuous 𝑑-norm𝑇 in (π‘Ž, 1) for all 0 < π‘Ž < 1; 0 < πœƒ < 1; 𝛿 ∈ βˆ† and 𝑓, 𝑔: 𝑄 β†’ 𝑃, we have (i). 𝛿(𝐹𝑓𝛼,𝑓𝛽(πœƒπ‘₯), 𝐹𝑔𝛼,𝑔𝛽(π‘₯), 𝐹𝑓𝛼,𝑔𝛼(π‘₯), 𝐹𝑓𝛽,𝑔𝛽(πœƒπ‘₯)) β‰₯ 0 for every 𝛼, 𝛽 ∈ 𝑄 and positive values of π‘₯ (ii). 𝑓(𝑄) βŠ‚ 𝑔(𝑄) (iii).𝛼0, 𝛼1 ∈ 𝑄 so that 𝑓𝛼0 = 𝑔𝛼1 and lim π‘›β†’βˆž 𝑇𝑖 (𝐹𝑓𝛼0,𝑓𝛼1 (π‘Ÿπ‘–)) = 1for π‘Ÿ > 1 and 𝑖 = 𝑛 π‘‘π‘œ ∞ Then, we can say to have common coincidence value for𝑓and 𝑔. Proof: Suppose, 𝛼0, 𝛼1 ∈ 𝑄 so that 𝑓𝛼0 = 𝑔𝛼1 and lim π‘›β†’βˆž 𝑇𝑖 (𝐹𝑓𝛼0,𝑓𝛼1 (π‘Ÿπ‘–)) = 1 (1) As 𝑓(𝑄) βŠ‚ 𝑔(𝑄) and 𝑓𝛼0 = 𝑔𝛼1, so a sequence < 𝛼𝑛 > can be constructed to have 𝑓𝛼𝑛 = 𝑔𝛼𝑛+1 Consider,𝑧𝑛 = 𝑓𝛼𝑛 Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 9s (2025) 2761 https://internationalpubls.com Then, (1) gives, lim π‘›β†’βˆž 𝑇𝑖 (𝐹𝑧0,𝑧1 (π‘Ÿπ‘–)) = 1 ; 𝑖 = 𝑛 π‘‘π‘œ ∞. Substituting 𝛼 = 𝛼𝑛 and 𝛽 = 𝛽𝑛+1 in (𝑖), we get 𝛿(𝐹𝑓𝛼𝑛,𝑓𝛼𝑛+1 (πœƒπ‘₯), 𝐹𝑔𝛼𝑛,𝑔𝛼𝑛+1 (π‘₯), 𝐹𝑓𝛼𝑛,𝑔𝛼𝑛 (π‘₯), 𝐹𝑓𝛼𝑛+1,𝑔𝛼𝑛+1 (πœƒπ‘₯)) β‰₯ 0, 𝑖. 𝑒. 𝛿(𝐹𝑧𝑛,𝑧𝑛+1 (πœƒπ‘₯), πΉπ‘§π‘›βˆ’1,𝑧𝑛 (π‘₯), πΉπ‘§π‘›βˆ’1,𝑧𝑛 (π‘₯), 𝐹𝑧𝑛,𝑧𝑛+1 (πœƒπ‘₯)) β‰₯ 0 and using virtue of 𝛿, for all 𝑛, we obtain 𝐹𝑧𝑛,𝑧𝑛+1 (πœƒπ‘₯) β‰₯ πΉπ‘§π‘›βˆ’1,𝑧𝑛 (π‘₯) So, by above Lemma, the sequence < 𝑧𝑛 >= < 𝑓𝛼𝑛 > is Cauchy. If 𝑔(𝑄) is 𝐹- complete, then there must be 𝛼 ∈ 𝑔(𝑄), so that,< 𝑧𝑛 > β†’ 𝛼, and for any 𝑧 ∈ 𝑄 so that 𝑔𝑧 = 𝛼. Substituting 𝛼 = 𝛼𝑛 and 𝑦 = 𝑧, in (𝑖), we have 𝛿(𝐹𝑓𝛼𝑛,𝑓𝑧(πœƒπ‘₯), 𝐹𝑔𝛼𝑛,𝑔𝑧(π‘₯), 𝐹𝑓𝛼𝑛,𝑔𝛼𝑛 (π‘₯), 𝐹𝑓𝑧,𝑔𝑧(πœƒπ‘₯)) β‰₯ 0 𝑖. 𝑒. 𝛿(𝐹𝑧𝑛,𝑓𝑧(πœƒπ‘₯), πΉπ‘§π‘›βˆ’1,𝑔𝑧(π‘₯), πΉπ‘§π‘›βˆ’1,𝑓𝑧(π‘₯), 𝐹𝑔𝑧,𝑓𝑧(πœƒπ‘₯)) β‰₯ 0 As 𝑛 β†’ ∞, we get, 𝛿(𝐹𝑔𝑧,𝑓𝑧(πœƒπ‘₯), 𝐹𝑔𝑧,𝑔𝑧(π‘₯), 𝐹𝑔𝑧,𝑓𝑧(π‘₯), 𝐹𝑓𝑧,𝑔𝑧(πœƒπ‘₯)) β‰₯ 0 𝑖. 𝑒. 𝛿(𝐹𝑔𝑧,𝑓𝑧(πœƒπ‘₯), 1, 𝐹𝑔𝑧,𝑓𝑧(π‘₯), 𝐹𝑓𝑧,𝑔𝑧(πœƒπ‘₯)) β‰₯ 0 Again, as 𝐹𝑔𝑧,𝑓𝑧(πœƒπ‘₯) β‰₯ 1, so 𝑓𝑧 = 𝑔𝑧 = 𝛼 Taking limit as nβ†’ο‚₯ , we have, Hence, 𝑓 and 𝑔 have𝑧 as a coincidence point and thus, for 𝑓(𝑄) to be complete, we have < 𝑧𝑛 > β†’ 𝛼 ∈ 𝑓(𝑄) βŠ‚ 𝑔(𝑄) so, 𝑧 is a common coincidence point. Theorem 4.2: With (𝑃, 𝐹, 𝑇)a generalized PM-space and a continuous 𝑑-norm𝑇 in (π‘Ž, 1) for all 0 < π‘Ž < 1; 0 < πœƒ < 1; 𝛿 ∈ βˆ† and 𝑓, 𝑔: 𝑄 β†’ 𝑃, we have (i). 𝛿(𝐹𝑓𝛼,𝑓𝛽(πœƒπ‘₯), 𝐹𝑔𝛼,𝑔𝛽(π‘₯), 𝐹𝑓𝛼,𝑔𝛼(π‘₯), 𝐹𝑓𝛽,𝑔𝛽(πœƒπ‘₯)) β‰₯ 0 for every 𝛼, 𝛽 ∈ 𝑄 and positive values of π‘₯ (ii). 𝑓(𝑃) βŠ‚ 𝑔(𝑃) (iii). βˆƒ 𝛼0, 𝛼1 so that 𝑓𝛼0 = 𝑔𝛼1 and lim π‘›β†’βˆž 𝑇𝑖 (𝐹𝑓𝛼0,𝑓𝛼1 (π‘Ÿπ‘–)) = 1for π‘Ÿ > 1 and 𝑖 = 𝑛 π‘‘π‘œ ∞ (iv). Either 𝑓(𝑃) or 𝑔(𝑃) is 𝐹- complete (v). 𝑓 and 𝑔 are commuting functions that commute at their coincidence point. Then, 𝑓 and 𝑔 have a common invariant point that is unique in, 𝑃. Proof: Let us choose 𝑃 = 𝑄 in previous theorem, then we obtain 𝑧𝑛 = 𝑓𝛼𝑛 and sequence < 𝑧𝑛 > is Cauchy. If 𝑔(𝑃) is 𝐹- complete, then 𝑧𝑛 β†’ 𝛼 ∈ 𝑔(𝑃), and we have 𝑧 ∈ 𝑃 so that 𝑔(𝑧) = 𝛼. Replacing 𝛼 by 𝛼𝑛 and 𝛽 by 𝑧 in inequality(𝑖) we have; 𝑓𝑧 = 𝑔𝑧 = 𝛼. As 𝑓 and 𝑔 are commuting, therefore, we must have, 𝑓𝑔𝑧 = 𝑔𝑓𝑧, so that,𝑓𝛼 = 𝑔𝛼. Again, using 𝛼 = 𝑧 and 𝛽 = 𝑓𝑧 in inequality (𝑖), we get 𝛿(𝐹𝑓𝑧,𝑓𝑓𝑧(πœƒπ‘₯), 𝐹𝑔𝑧,𝑔𝑓𝑧(π‘₯), 𝐹𝑔𝑧,𝑓𝑧(π‘₯), 𝐹𝑓𝑓𝑧,𝑔𝑓𝑧(πœƒπ‘₯)) β‰₯ 0 Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 9s (2025) 2762 https://internationalpubls.com 𝑖. 𝑒. 𝛿(𝐹𝛼,𝑓𝛼(πœƒπ‘₯), 𝐹𝛼,𝑓𝛼(π‘₯), 1, 𝐹𝑓𝛼,𝑔𝛼(πœƒπ‘₯)) β‰₯ 0 Using the properties of distance metric(𝛿) for all positive,π‘₯ we obtain, 𝐹𝛼,𝑓𝛼(πœƒπ‘₯) β‰₯ 𝐹𝛼,𝑓𝛼(π‘₯) which is true only if 𝛼 = 𝑓𝛼, otherwise we reach to a contradiction with respect to Lemma 4.1. This implies that, 𝛼 is a common invariant point of 𝑓 and 𝑔. To prove the uniqueness, we suppose 𝛼 and 𝛽 are two common invariant points of 𝑓 and 𝑔. Using (𝑖), we get 𝛿(𝐹𝛼,𝛽(πœƒπ‘₯), 𝐹𝛼,𝛽(π‘₯), 𝐹𝛼,𝛼(π‘₯), 𝐹𝛽,𝛽(πœƒπ‘₯)) β‰₯ 0β‡’ 𝐹𝛼,𝛽(πœƒπ‘₯) β‰₯ 1 β‡’ 𝛼 = 𝛽. This shows the uniqueness of the result. Definition 4.1: [6] Aforementioned t–norm T is said HadΕΎiΔ‡ type or 𝐻- type, denoted as 𝑇 ∈ 𝐻, if the group {𝑇𝑛}π‘›βˆˆπ‘ of it iterates is inductively depicted, βˆ€π‘₯ ∈ [0,1] by, 𝑇0(π‘₯) = 1, 𝑇𝑛+1(π‘₯) = 𝑇(𝑇𝑛(π‘₯), π‘₯), for all non-negative π‘₯ and as 𝑇𝑛 is defined, it is equi-continuous at π‘₯ = 1. This states the following, νœ€ ∈ (0,1)βˆƒ a πœ— ∈ (0,1), so that, π‘₯ > 1 βˆ’ πœ—β‡’ 𝑇𝑛(π‘₯) > 1 βˆ’ νœ€, for all 𝑛 β‰₯ 1. The statement is followed by HadΕΎiΔ‡: Statement 4.1: [6] (i). If we take 𝑇 β‰₯ 𝑇𝐿, then for 𝑖 = 𝑛 π‘‘π‘œ ∞, lim π‘›β†’βˆž 𝑇𝑖π‘₯𝑛+1 = 1, both sides imply βˆ‘ 1 βˆ’ π‘₯𝑛 < ∞ and 𝑇 = 𝑇¡ π‘†π‘Š (By 𝑇 β‰₯ 𝑇𝐿 𝑖. 𝑒. π‘Ž β‰₯ 𝑐, 𝑏 β‰₯ 𝑑⇒ 𝑇(π‘Ž, 𝑏) = 𝑇(𝑐, 𝑑)). (ii). When, 𝑇 ∈ 𝐻, then for each (π‘₯𝑛)π‘›βˆˆπ‘ in [0,1], lim π‘›β†’βˆž π‘₯𝑛 = 1β‡’ lim π‘›β†’βˆž 𝑇𝑖π‘₯𝑛+𝑖 = 1,for 𝑖 = 𝑛 π‘‘π‘œ ∞. (iii). When, 𝑇 ∈ π‘‡πœ‡ 𝐷, π‘‡πœ‡ 𝐴𝐴𝑇, then for 𝑖 = 𝑛 π‘‘π‘œ ∞, lim π‘›β†’βˆž 𝑇𝑖π‘₯𝑛+𝑖 = 1, both sides imply βˆ‘(1 βˆ’ π‘₯𝑛)πœ‡ < ∞. We wish to repeat the notion of distribution functions and PMS for the interval[0, ]ο‚₯ . Definition 4.2: [6] Consider a class of distribution map𝐿+as 𝐹: [0, ∞] β†’ [0, ∞] with, (i). 𝐹 has zero value at π‘₯ = 0 (ii). 𝐹 is a non-decreasing map and (iii). 𝐹 is left continuous for 0 < π‘₯ < ∞. Let 𝐷+ βŠ‚ 𝐿+ contains𝐹, so that, lim π‘›β†’βˆž 𝐹(π‘₯) = 1. Then, the distribution function, νœ€0 is defined as νœ€0 νœ€0 = { 0, π‘₯ = 0 1, π‘₯ > 0 lies in 𝐷+. Statement 4.2: Consider (𝑃, 𝐹, 𝑇) asa generalized PM-space and a continuous 𝑑-norm𝑇 ∈ 𝐻 with 0 < πœƒ < 1; 𝛿 ∈ βˆ† and 𝑓, 𝑔: 𝑃 β†’ 𝑃, so that, (i). 𝛿(𝐹𝑓𝛼,𝑓𝛽(πœƒπ‘₯), 𝐹𝑔𝛼,𝑔𝛽(π‘₯), 𝐹𝑓𝛼,𝑔𝛼(π‘₯), 𝐹𝑓𝛽,𝑔𝛽(πœƒπ‘₯)) β‰₯ 0 for every 𝛼, 𝛽 ∈ 𝑃 and positive values of π‘₯ (ii). 𝑓(𝑃) βŠ‚ 𝑔(𝑃) (iii). βˆƒ 𝛼0, 𝛼1 so that 𝑓𝛼0 = 𝑔𝛼1 for which 𝐹𝑓𝛼0,𝑓𝛼1 ∈ 𝐷+ (iv). Either 𝑓(𝑃) or 𝑔(𝑃) is 𝐹- complete Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 9s (2025) 2763 https://internationalpubls.com (vi). 𝑓 and 𝑔 are commuting functions that commute at their coincidence point. Then, 𝑓 and 𝑔 have a unique common invariant point in 𝑃. Proof: In Theorem 4.1, if we consider a πœ‡ > 1 with a sequence < 𝑧𝑛 > along with a mapping, 𝑓𝑧0,𝑧1 ∈ 𝐷+, so that, lim π‘›β†’βˆž 𝐹𝑧0,𝑧1 (πœ‡π‘›) = 1. Therefore, using (𝑖𝑖) of Statement 4.1, we get lim π‘›β†’βˆž 𝑇𝑖 (𝐹𝑧0,𝑧1 (πœ‡π‘–)) = 1; 𝑖 = 𝑛 π‘‘π‘œ ∞. Hence, Theorem 4.2 follows the required results. 5. Discussion Probabilistic spaces offer a powerful framework for handling the uncertainties and variability inherent in real-world data [12]. By incorporating probability distributions into the metric space, we can model the stochastic nature of data more effectively. The Jungck contraction theorem provides a robust tool for analysing the convergence of iterative processes in probabilistic metric spaces, which is crucial for the stability and reliability of ML algorithms. This work has demonstrated the application of probabilistic metric spaces and the Jungck contraction theorem in enhancing the robustness and convergence properties of ML algorithms. Through case studies on probabilistic K-means clustering and probabilistic neural networks, we have illustrated the practical benefits of this approach. Future research can extend these concepts to other ML algorithms and explore the integration of probabilistic metric spaces with other probabilistic modelling techniques. In a nutshell, the probabilistic metric space approach, underpinned by the Jungck contraction theorem, offers a promising direction for developing more robust and reliable ML algorithms in the face of uncertainty and variability. References [1] S. Banach, β€œSur les opΓ©rations dans les ensembles abstraits et leur application aux Γ©quationsintΓ©grales,” Fundamenta Mathematicae, 3(1), pp.133-181,1922. [2] K. Menger, β€œStatistical metrics,” Proceedings of the National Academy of Sciences of the United States of America, 28(12), pp.535-537, 1942. [3] E. Fix, and J. L. Hodges, β€œDiscriminatory analysis-nonparametric discrimination: Consistency properties. Report Number 4, Project 21-49-004,” USAF School of Aviation Medicine, Randolph Field, Texas,1951. [4] G. Jungck, β€œCompatible mappings and common fixed points,” International Journal of Mathematics and Mathematical Sciences, 1(4), pp.215-222,1976. [5] C. C. Y. Dorea, β€œOn probabilistic metric spaces,” Bulletin of the Brazilian Mathematical Society, 14(2), pp.99-118,1983. [6] O. HadΕΎiΔ‡ and E. Pap, β€œProbabilistic metric spaces. In: Fixed point theory in probabilistic metric spaces,” Mathematics and its Appl. Vol. 536, Springer, Dordrecht, pp.47-94, 2001.https://doi.org/10.1007/978-94-017-1560-7_2 [7] O. Banerjee, S. Merugu, I. S. Dhillon, and J. Ghosh, "Clustering with Bregman Divergences." Journal of Machine Learning Research, vol. 6, pp. 1705-1749, 2005. [8] T. Hastie, R. Tibshirani, J. Friedman, β€œThe Elements of Statistical Learning: Data Mining, Inference, and Prediction (2nd ed.),” Springer, 2009. Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 9s (2025) 2764 https://internationalpubls.com [9] I. Goodfellow, Y. Bengio, & A. Courville, β€œDeep Learning,” MIT Press, 2016. [10] R. S. Suttonand A. G. Barto, β€œReinforcement Learning: An Introduction (2nd ed.)”. MIT Press, 2018. [11] A. Mishra, P. K. Tripathi, A. K. Agrawal, D. R. Joshi, β€œA Probabilistic Approach in Modeling the Growth of Cell Population,” Test Engineering and Management, Vol. 83, pp. 737-742, 2020. [12] A. Mishra, Achieving Environmental Sustainability Through Contraction Mapping in Growth of Cell Population, International Journal of Engineering Research and Management Technology, ISSN:2348-403, International Conference on Mathematical Sciences and Computational Intelligence (ICMSCI-2020), (2020), 35-36. [13] V. Fortuin, O. E. Ganeaand M. Welling, β€œDeep generative models with learnable knowledge priors,” In International Conference on Machine Learning (ICML), 2021. [14] A. Mishra and P. K. Tripathi, β€œIteration-based reduction in cell population for biomedical applications”, Proceedings of Advances in AI for Biomedical Instrumentation, Electronics and Computing, Taylor and Francis Group, CRC Press, UK, pp.466-72, 2024. http://dx.doi.org/10.1201/9781032644752-86 [15] A. Mishra and P. K. Tripathi, β€œApplicability of Intimate Mapping in Digital Image Sources”, Grenze International Journal of Engineering and Technology (Grenze Scientific Society), 10 (1), 711-715, 2024. [16] A. Mishra and P. K. Tripathi, β€œIterative Methods under Jungck Contraction Theorem for Root Finding in Numerical Analysis for AI Applications”, Grenze International Journal of Engineering and Technology (Grenze Scientific Society), 11 (1), 4703-4708, 2025. http://dx.doi.org/10.1201/9781032644752-86