Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 3 (2025) 488 https://internationalpubls.com Trends and Growth of Productivity of Agricultural Sector using Statistics Through Machine Learning in Assam Smrita Borthakur1, Ranjan Kr. Sahoo2, Sarat Ch. Kakaty3, Supahi Mahanta* 1Research Scholar, Department of Statistics, Central University of Haryana, India ORCID-0000-0001-9734-3795 2Department of Statistics, Central University of Haryana, India 3Department of Statistics, Dibrugarh University Assam, India *Department of Statistics, Assam Agricultural University Assam, India ORCID-0000-0002-7200-4436 *Corresponding author. E. mail: supahi.mahanta@aau.ac.in Article History: Received: 02-08-2024 Revised: 23-09-2024 Accepted: 05-10-2024 Abstract: The study of trend and growth of agricultural sector by applying statistical techniques through machine learning analysis has become a powerful tool in in the long run. period, In this study Trend Analysis have been done using fitting of curves of the types, Exponential, Modified Exponential, Gompertz and Power Curves by taking secondary data for a period of 1993to 2023. The data have been collected from the official records of Directorate of Economics and Statistics Assam. Trend analysis have been done using Mann-Kendall test and Sen’s slope estimator. From the growth models modified exponential is adjudged by higher 𝑅2 and lowest MSE and lowest AIC with significant t values. Keywords: Trend, Growth, Exponential, Modified Exponential, Gompertz and Power Curves, Mann- Kendall test and Sen’s slope estimator 1. Introduction Growth is a key process for development. Kumar et al. (2017), Fugile (2018), in their study talked on the deceleration of productivity, highlighting its implications for global food security. Sherman et al. (1950) together with Spurr and Arnold (1948) are the pioneers for showing Simplified Procedures for Fitting Growth Curves for a given set of data. In the realm of agricultural economics, several seminal studies have applied a diverse range of growth curve models to unravel trends in agricultural development. (Aheam et al. (1998), Greeshma et al. (2014), Jaypatre et al. (2010), Sreenivas et al. (2019), Suthar et al. (2024)) conducted pioneering research, employing Exponential, Modified Exponential, Logistic, Gompertz, and Power Curves to analyze agricultural trends comprehensively. Their work not only involved the computation of growth rates and compound growth rates but also provided valuable insights into the dynamics of agricultural area, production, and productivity in terms of GDP. This foundational research laid the groundwork for subsequent investigations, such as Adinata et al. (2022) and Brand et al. (2024), which further delved into the intricacies of agricultural growth for animals using similar methodologies. mailto:supahi.mahanta@aau.ac.in Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 3 (2025) 489 https://internationalpubls.com 𝑦 𝑑 Building upon these early works, Williams et al. (2003) explored the implications of growth curve analysis for agricultural policy formulation. Subsequent studies by Rahman et al. (2011) and Li et al. (2010) expanded the scope of inquiry by examining the role of technological advancements in driving agricultural productivity growth, utilizing sophisticated modelling techniques. In parallel, Mann-Kendall and Sen's slope analysis has emerged as a powerful tool for trend analysis in agricultural economics. Studies by (Binod et al. (2022), Ham et al. (2023) and Chandler et al. (2011)) have utilized this method to discern long-term patterns in agricultural production and productivity, reinforcing traditional growth curve analyses. Chen et al. (2022) suggested applications machine learning techniques for trend analysis. Furthermore, the integration of exploratory data analysis (EDA) techniques, as demonstrated by Shanmugan et al (2018) and Alston et al. (2019), has enriched our understanding of the underlying drivers of agricultural growth, facilitating the identification of key factors influencing agricultural dynamics. These studies collectively contribute to a deeper understanding of the complex dynamics shaping agricultural economies worldwide, highlighting the pivotal role of growth curve analysis, Mann-Kendall and Sen's slope analysis, and exploratory data analysis in elucidating trends and informing policy decisions. 2. Data and methodology: Data collection: The time-series data on fish productivity (in Kg/hectare) in Assam for a period of 30 years from 1993 to 2023 was obtained from the Directorate of Economics and Statistics, Government of Assam for conducting the present study. The methods and Models we have used here are least squares, Men Kandell Sen’s Slope, 3-point sum using Exploratory data analysis in python. The models that have been employed are (A) Exponential, logistic, Modified Exponential, Gompertz and Power Curve β€’ 𝑦𝑑=a 𝑏𝑑 (1) Growth rate of the equation is 𝑑𝑦𝑑 = a 𝑑𝑏𝑑 𝑑𝑑 𝑑𝑦𝑑 𝑑𝑑 = a log bβˆ— 𝑏𝑑 =log b βˆ— 𝑦𝑑 𝑑𝑑⁄ = log b (2) 𝑑 β€’ 𝑦 =a𝛽𝑑𝛾𝑑 2 (3) Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 3 (2025) 490 https://internationalpubls.com = a 𝑑 𝑑 𝑑 𝑑 𝑦 𝑦𝑑 Growth rate of the above equation is 𝑑𝑦 2 𝑑 𝛽𝑑 𝑑𝛾𝑑 2 𝑑= a[𝛾𝑑 + 𝛽𝑑 ] 𝑑𝑑 𝑑 𝑑 𝑑𝑑 2 2 [𝛾 π‘™π‘œπ‘”π›½ 𝛽 + 𝛽 2π‘™π‘œπ‘”π›Ύ 𝛾 ] 𝑑 𝑦𝑑 = Y logΞ²+ Y2log𝛾 𝑑𝑑⁄ = (logΞ²+2log𝛾) (4) 𝑑 β€’ 𝑦𝑑=K +a𝑏𝑑 (5) Growth rate of the above equation is 𝑑𝑦𝑑 = =alogb𝑏𝑑 𝑑𝑑 𝑑 𝑦 𝑑 =log b a𝑏𝑑 𝑑𝑑⁄ =π‘™π‘œπ‘”π‘(1-K/𝑦𝑑) (6) Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 3 (2025) 491 https://internationalpubls.com 𝑦 𝑦 β€’ 𝑦𝑑=K a𝑏𝑑 (7) Growth rate of the above equation is 𝑑𝑦𝑑 =K 𝑑𝑏𝑑 𝑑𝑑 𝑑𝑦𝑑 𝑑𝑑 =K log a a𝑏𝑑 =log a 𝑦𝑑 𝑑𝑑⁄ = log a (8) 𝑑 β€’ 𝑦𝑑=A π‘’π‘Ÿπ‘‘ (9) Growth rate of the above equation is 𝑑𝑦𝑑 𝑑𝑑 = A r π‘’π‘Ÿπ‘‘ =r 𝑦𝑑 𝑑𝑦𝑑 𝑑𝑑⁄ = r (10) 𝑑 β€’ 𝑦𝑑=𝛼𝑑𝛽 (11) Growth rate of the above equation is 𝑑𝑦𝑑 = 𝛼 𝑑𝑑𝛽 𝑑𝑑 𝑑𝑑 = π›ΌΞ²π‘‘π›½βˆ’1 𝑑𝑦𝑑 𝑑𝑑⁄ = 𝛽 (12) 𝑦𝑑 𝑑 To determine the magnitude and direction of the trend Mann Kendall and Sen’s slope have been used. (B) Mann-Kendall Test: Rank-based nonparametric methods provide alternative statistical approaches to the conventional parametric methodology. Recent advancements in the field of individual rank tests have been substantially influenced by the works of Zimmerman and Wolfowitz (1940), Hotelling and Pabst (1936), Kendall (1938), Smirnov (1939), Mann (1945), and Wald and Wolfowitz (1940), despite the older nature of the concept. Savage came across over three thousand publications pertaining to nonparametric procedures, Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 3 (2025) 492 https://internationalpubls.com particularly rank-based methods, in 1962. Prior to these advancements, a vast quantity of research papers were published concerning nonparametric methods. However, during that period, studies were concerned about the efficacy of nonparametric methods due to the weak assumptions they required for validity. Historiographic validation of the Wilcoxon test, along with other nonparametric techniques, is attributable to Chernoff and Savage (1958) and Hodge and Lehmann (1956), provided that the normality assumption is satisfied. However, in circumstances where the normality assumption is not met, the procedures might prove to be more impactful and legitimate. Mann-Kendall and Modified Mann-Kendall (Haled and Rao, 1998) are the primary non-parametric techniques utilized in hydroclimatic time series trend analysis (Kendall, 1955; Mann, 1945). With regard to the orientation of the trend, the Mann-Kendall's trend test examines time series data for trends, notwithstanding its non-parametric nature (Kisi et al. 2004, Partal and Kahya, 2006; Yenigun et al., 2008; Hadgu et al., 2013, Alhaji et al. (2018)). Consequently, its vulnerability to anomalous effects is diminished. The MK test compares the alternative hypothesis, which asserts the presence of a trend in either an upward or downward direction, to the null hypothesis, which states that no trend exists. The MK test statistic (S) is given by- Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 3 (2025) 493 https://internationalpubls.com nβˆ’1 n S=οƒ₯οƒ₯ sgn(xj βˆ’ xk ) k =1 j=k +1 +1 if (π‘₯𝑗 βˆ’ π‘₯π‘˜)>0 sgn (π‘₯𝑗 βˆ’ π‘₯π‘˜) = 0 if (π‘₯𝑗 βˆ’ π‘₯π‘˜)=0 (13) -1 if (π‘₯𝑗 βˆ’ π‘₯π‘˜)<0 Var=[n(n βˆ’ 1)(2n + 5) βˆ’ βˆ‘ t(t βˆ’ 1)(2t + 5)]/18 (14) The notation βˆ‘t represents the summation of all ties, while t represents the extent of a specific tie. In cases where the sample size surpasses 10, the standard normal variate z is computed using Equation (2) (Douglas (15) et al., 2000). (S – 1)/√Var(S) if S>0 z = 0 if S=0 (S + 1)/√Var(S) if S<0 Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 3 (2025) 494 https://internationalpubls.com ⁄ i # To import the data For a two-sided test, H0 should be accepted if z≀ zΞ±/2 at the given level of significance. A positive value of S indicates an upward trend and a negative value of S indicates downward trend. (C) Sen’s Slope: The magnitude of trend in a time series is generally determined by using regression analysis and Sen’s estimator (Sen, 1968). The former one is a parametric test and the later one is non-parametric test. Both the two methods assume a linear trend in the time series. In this method, the slopes (Ti) for all data pairs are first calculated by the following way- T = Xjβˆ’Xk jβˆ’k for i= 1,2, …., N (16) here xj are the values at time j and xk are the values at time k (j > k). The median of these N values of Ti is Sen’s estimator of slope, which is calculated as TN+1 if N is odd 2 L= (17) 1 (TN + TN+2 ) if N is even 2 ⁄2 ⁄2 A time series is characterized by an upward trend when the value of L is positive, and a downward trend when the value is negative. Checking of trend of productivity of crops was done by trend package included in the software Python (D) EDA: EDA is a part of ML technique used to analyze data using both non-visual and visual technique. Tukey (1977) and Mathew et al. (2016) mentioned the graphical and non-graphical methods of Exploratory data analysis for applied sciences for trend analysis. There is a structured approach which is followed Involves thorough analysis of data to understand the current situation. Through EDA we can extract meaningful information based on domain understanding. The codes involved in EDA are Importing the dataset Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 3 (2025) 495 https://internationalpubls.com def remove_outlier(col): sorted(col) Q1, Q3= np. Percentile (col, [25,75]) IQR=Q3-Q1 lower_range= Q1-(1.5 * IQR) upper_range= Q3+(1.5 * IQR) return lower_range, upper_range for column in df. Iloc [ :, 1:7]. columns: lr, ur=remove_outlier(df[column]) df[column]=np. where (df[column]>ur, ur, df[column]) df[column]=np. where (df[column]