American Scientific Research Journal for Engineering, Technology, and Sciences (ASRJETS) ISSN (Print) 2313-4410, ISSN (Online) 2313-4402 ยฉ Global Society of Scientific Research and Researchers http://asrjetsjournal.org/ On the Modification of M-out-of-N Bootstrap Method for Heavy-Tailed Distributions Hannah F. Opayinkaa*, Adedayo A. Adepoju b aFederal College of Education (Special),Nigeria, Phone:+2348034975272 b Statistics Department, University of Ibadan, Nigeria, Phone:+2348066430258 afolashadeopayinka@gmail.com bpojuday@yahoo.com Abstract This paper is on the modification of ๐‘š-out-of-๐‘› bootstrap method for heavy-tailed distributions such as income distribution. The objective of this paper is to present a modified ๐‘š-out-of-๐‘› bootstrap method (๐‘š๐‘š๐‘œ๐‘œ๐‘›) and compare its performance with the existing m-out-of-n bootstrap method (๐‘š๐‘œ๐‘œ๐‘›). The nature of the upper tail of a distribution is the major reason for the poor performance of classical bootstrap methods even in large samples. The โ€˜๐‘š๐‘š๐‘œ๐‘œ๐‘›โ€™ bootstrap method was therefore, proposed as an alternative method to โ€˜๐‘š๐‘œ๐‘œ๐‘›โ€™ bootstrap method. The distribution involved has finite variance. The simulated data sets used was drawn from Singh-Maddala distribution. The methodology involved decomposing the empirical distribution and sampling only nโƒ› times with replacement from a sample size n, such that nโƒ› โ†’ โˆž as n โ†’ โˆž, and nโƒ›/n โ†’ 0. The performances are judged using standard error; absolute bias; coefficient of variation and root mean square error. The findings showed that โ€˜๐‘š๐‘š๐‘œ๐‘œ๐‘›โ€™ performed better than ๐‘š๐‘œ๐‘œ๐‘› in moderate and larger samples and it converged faster. Keywords: Bootstrap; Decomposition; Heavy-tailed distributions; Singh-Maddala distribution. 1. Introduction 1.1. Background of the Study This study is concerned with finding a reliable alternative bootstrap method to heavy-tailed distributions. The tail of a distribution, especially the upper tail affects the performance of bootstrap methods. ------------------------------------------------------------------------ * Corresponding author. E-mail address: folashadeopayinka@gmail.com 142 http://asrjetsjournal.org/ American Scientific Research Journal for Engineering, Technology, and Sciences (ASRJETS) (2015) Volume 14, No 1, pp 142-155 Bootstrap is a resampling procedure developed by Efron in 1979. The usual practice when estimating the properties of an estimator (such as its variance) is by measuring those properties when sampling from an approximating distribution. Empirical distribution of the observed data is a standard choice for an approximating distribution. The resampling could be done from an independent and identically distributed (๐‘–๐‘–๐‘‘) population. Bootstrap is used when parametric assumptions are in doubt, when the formulas for the calculation of standard errors are complicated in parametric inference. There is no need to force Gaussian or any other parametric distributional assumptions on data. The distribution could be skewed, multimodal, and heavy-tailed; the estimator of interest could be complicated [1]. The use of bootstrap has been applied to many estimators within cross-section data [2]. In heavy-tailed distributions like income distribution, there is frequency of outliers in data sets, which usually cause difficulties in the use of asymptotic and bootstrap methods. The nature of the upper-tail of the distribution generally affects the performance of the methods [3]. The author in [4] proposed the use of the bootstrap for the most commonly applied procedures in inequality, mobility and poverty measurements. He suggests that the simplest possible bootstrap procedure should be the preferred method in practice because it achieves precision and takes into account the stochastic dependencies in the data, without the need of dealing with its covariance structure explicitly. His simulation results suggest that the bootstrap performs well in finite samples. He also decomposed income distribution into subgroup such that; ๐‘ค๐‘–, ๐‘ฅ๐‘– are the weights and income sources respectively. His study regarded all subgroups to constitute the population and all subgroups are disjoint. The authors in [3] studied finite-sample performance of asymptotic and bootstrap inference for both inequality and poverty measures. Their simulation results showed that neither asymptotic nor classical bootstrap inference for inequality measures perform well, even when the sample size is large enough. They found that the performances of both asymptotic and bootstrap are affected by the nature of the upper-tail of the income distribution. Authors of some studies [4, 5] involving heavy-tailed distribution recommend the use of bootstrap rather than asymptotic methods. After the publication of Efron in 1982, research activity on the bootstrap grew so fast with the emergence of many theoretical developments on the asymptotic consistency of bootstrap estimate coupled with real-world applications. Focus changed in 1990s to finding applications and variants that would perform well in practice. Some studies on bootstrap inference for inequality measures were done and the use of bootstrap methods rather than asymptotic methods was recommended [1]. Heavy-tail means that the probability of getting very large values is high. Therefore, heavy-tailed distributions typically represent wild as opposed to mild randomness, examples are income distributions, financial returns, insurance payouts, reference links on the web etc. A technical difficulty is that, all moments do not exist for these distributions. Heavy-tailed distributions that are often used include: Burr (Singh-Maddala) distribution, Pareto distribution, 143 American Scientific Research Journal for Engineering, Technology, and Sciences (ASRJETS) (2015) Volume 14, No 1, pp 142-155 L๏ฟฝฬ๏ฟฝvy distribution, Weibul distribution, Log-gamma distribution, etc. Singh-Maddala (Burr) distribution is a member of a system of continuous distributions introduced by Burr in 1942.The Burr distribution/ Singh- Maddala distribution is a continuous probability distribution for a non-negative random variable. It is most commonly used to model household income [6]. This study proposes an alternative bootstrap method called โ€˜๐‘š๐‘š๐‘œ๐‘œ๐‘›โ€™ which is expected to perform better than ๐‘š๐‘œ๐‘œ๐‘›. It has been established that ๐‘š๐‘œ๐‘œ๐‘› could overcome the inconsistency in the classical method. The study would involve estimating the income data by obtaining the chosen estimator in each bootstrap method, comparing the statistical inference of ๐‘š๐‘œ๐‘œ๐‘› and ๐‘š๐‘š๐‘œ๐‘œ๐‘› bootstrap methods. Using simulation would allow one to assess the reliability of these methods for empirical work. 1.2. Limitation of the Study The research was carried out using 64-bit Operating system laptop computer. The work would have been faster if it had been done on a macro computer. 2. Methodology The research methodology is described as follows. 2.1. ๐’Ž-out ofโ€“๐’ (๐’Ž๐’๐’๐’) bootstrap method Authors in [3] regard moon bootstrap method as being useful when the classical bootstrap fails or when it is difficult to check its consistency. The author in [7] described ๐’Ž๐’๐’๐’ as sampling ๐’Ž times without replacement from a sample size ๐’, instead of sampling ๐’ times; such that ๐’Ž is much less than ๐’. Usually the asymptotic theory requires ๐’Ž โ†’ โˆž ๐š๐ฌ ๐’ โ†’ โˆž, but at a slower rate such that ๐’Ž/๐’ โ†’ ๐ŸŽ. In [8] it was called โ€™the great ๐’Ž ๐จ๐ฎ๐ญ ๐’ bootstrap with (๐’Ž/ ๐’ โ†’ ๐ŸŽ )โ€™ where the bootstrap sample size ๐ฆ is much smaller than the original size. Mathematically, the requirement is ๐’Ž โ†’ โˆž and ๐’Ž/๐’ โ†’ ๐ŸŽ,๐’‚๐’” ๐’ โ†’ โˆž. In theory, the problem is fixed, but in practice, some troubles are involved such as how to choose ๐’Ž. An obvious suggestion is to settle for a fraction of, say 20%. It was pointed out that in good situations, where the regular bootstrap performs, such a ๐’Ž is not advisable, it could result in loss of efficiency. Authors in [9] proposed the ๐ฆ-out-of-๐ง bootstrap with or without replacement, where m โ†’ โˆž and ๐ฆ ๐งโ„ โ†’ 0 as a way of ensuring consistency when the classical bootstrap is not consistent. Authors in [10] explored ๐ฆ-out-of-๐ง bootstrapping from the empirical distribution function in nonstandard problems and proved the consistency of the method. 2.2. modified ๐’Ž-out-of-๐’ (๐’Ž๐’Ž๐’๐’๐’) bootstrap method The ๐‘š๐‘š๐‘œ๐‘œ๐‘› bootstrap method is a modification of ๐‘š๐‘œ๐‘œ๐‘›, it involves decomposition of the empirical 144 American Scientific Research Journal for Engineering, Technology, and Sciences (ASRJETS) (2015) Volume 14, No 1, pp 142-155 distribution (๐น๏ฟฝ๐‘›) and sampling only nโƒ› times with replacement from a sample size n, such that nโƒ› โ†’ โˆž as n โ†’ โˆž, and nโƒ›/n โ†’ 0. The method would resample from ๐ป๐‘› distribution which is a decomposed version of ๐น๏ฟฝ๐‘›. The procedure provides and estimates different measures of statistical precision for an estimator ๐œƒ๏ฟฝ. The method satisfies the conditions in [11]. 2.3. Validity of (๐’Ž๐’Ž๐’๐’๐’) bootstrap method โ€ข Samples must be independently and identically distributed; [1] reported that ๐‘–๐‘–๐‘‘ works in large sample. โ€ข For a bootstrap approach to work well, it was suggested in [12] that the distribution function should have a differentiable density. This study makes use of simulated data sets drawn from Singh-Maddala distribution which [3] said can quite successfully mimic observed income distributions in various countries. It can be shown that the distribution has a differentiable density. The cumulative density function (CDF) of the distribution can be written as: ๐น(๐‘ฅ) = 1 โˆ’ 1 ๏ฟฝ1+๐‘Ž๐‘ฅ๐‘๏ฟฝ ๐‘ {1} Where a = scale parameter; b = shape parameter; c = shape parameter [3]. And the probability density function (pdf) as: ๐‘“(๐‘ฅ) = ๐‘Ž๐‘๐‘ ๐‘ฅ๐‘โˆ’1 ๏ฟฝ1+๐‘Ž๐‘ฅ๐‘๏ฟฝ ๐‘+1 {2} โ€ข The validity of bootstrap also requires that the estimator (a functional form of the empirical distribution function) converges to the true parameter value (the functional form for the true population distribution). The commonly used parameters of distribution function can be expressed as functional form of the distribution, which includes the mean, the variance etc. Sample estimates such as the sample mean can be expressed as the same functional form applied to the empirical distribution. This provides guideline for this method in choosing the mean as the functional form of the distribution. A functional form is simply a mapping that takes a function F into a real number, examples of such are the mean and variance of a distribution [3][13]. Assuming mean is used as the functional form in this study, let ๐œ‡ be the mean for a distribution function ๐น, then ๐œ‡ = โˆซ ๐‘ฅ๐‘‘๐น(๐‘ฅ). It can be shown that the functitonal form of the Singh-Maddala distribution exists: ๐ธ(๐‘ฅ) < โˆž ; ๐ธ(๐‘ฅ) = โˆซ ๐‘ฅ โˆ™ ๐‘๐‘‘๐‘“ ๐‘‘๐‘ฅ ๐œ‡ = ๐‘๐‘Ž 1 ๐‘ฮ“๏ฟฝ๐‘โˆ’1+1๏ฟฝฮ“(๐‘โˆ’๐‘โˆ’1) ฮ“(๐‘+1) [3] โ€ข It was shown in [11] that bootstrap principle works for sample mean when finite second moments exist. 145 American Scientific Research Journal for Engineering, Technology, and Sciences (ASRJETS) (2015) Volume 14, No 1, pp 142-155 This provided a stronger justification for this method. The second moment exists for the Singh-Maddala distribution considered in this study. ๐ธ(๐‘ฅ2) < โˆž ;๐ธ(๐‘ฅ2) = ๏ฟฝ ๐‘ฅ2๐‘“(๐‘ฅ)๐‘‘๐‘ฅ ๐ธ(๐‘ฅ2) = ๏ฟฝ ๐‘ฅ2 ๐‘Ž๐‘๐‘ ๐‘ฅ๐‘โˆ’1 (1 + ๐‘Ž๐‘ฅ๐‘)๐‘+1 ๐‘‘๐‘ฅ ; = ๐‘Ž๐‘๐‘ 2 โˆ’ ๐‘ [1 + ๐‘Ž๐‘ฅ๐‘]2โˆ’๐‘ โ€ข Bootstrapping actually works (i.e. consistent) if the following holds: If ๐‘‡(๐น๐‘›โˆ—) converges to ๐‘‡(๐น) as ๐‘› โ†’ โˆž, (i.e. the bootstrap estimate is consistent for the population parameter) which implies that โ€ข ๐‘‡(๐น๐‘›) converges to ๐‘‡(๐น) as ๐‘› โ†’ โˆž, (i.e sample estimate is consistent for the population parameter, when ๐น๐‘› converges to ๐น uniformly). โ€ข ๐‘‡๏ฟฝ๐น๏ฟฝ๐‘›โˆ—๏ฟฝ โˆ’ ๐‘‡(๐น๐‘›) โ†’ 0 as ๐‘› โ†’ โˆž, (i.e. the difference between bootstrap estimate and sample estimates tends to zero). The simulation study can also be used to confirm or deny the usefulness of the bootstrap estimate. The performances of the estimates are judged in the empirical work, by obtaining standard error, absolute bias, root mean square error and coefficient of variation [13]. 2.4. Decomposition of Empirical Distribution The proposed measurement scenarios are decomposition of the empirical distribution by sub-group (in form of strata) which could be income levels, the subgroups are disjoint and all subgroups taken together constitute the population. Authors in [5] included decomposition by population subgroups due to significant differences in income levels among individuals, these differences are caused by some characteristics such as age, race etc. The author in [1] recommended stratified sampling as a remedy for inconsistency in bootstrap method. Stratification can be useful in reducing the variability of some estimates. Stratified sampling has been regarded as a method of variance reduction in computational statistics and that a stratified survey could claim to be more representative of the population than a survey of simple random sampling or systematic sampling. In stratified sampling, there is assurance that estimates would be made with equal precision in different parts of the region, and that comparisons of sub-regions would be made with equal statistical power [6]. The author in [4] decomposed income distribution into subgroups such that ๐‘ค๐‘–, ๐‘ฅ๐‘– are the weights and the income sources. It is assumed that the subgroups are disjoint and that all subgroups taken together constitute the 146 American Scientific Research Journal for Engineering, Technology, and Sciences (ASRJETS) (2015) Volume 14, No 1, pp 142-155 population (i.e. are mutually exclusive).His statistics of interest are the contribution of income sources to overall inequality. The observed data is assumed to be of the form ๐‘ฆ๐‘– = (๐‘Ÿ๐‘– ,๐‘ฅ๐‘–) for ๐‘– = 1 โ€ฆ . . ๐‘›, which can be interpreted as ๐‘–๐‘–๐‘‘ sample of size ๐‘› from a joint distribution of ๐บ of ๐‘Ÿ and ๐‘ฅ, ๐บ(๐‘Ÿ๐‘– ,๐‘ฅ๐‘–), let ๐‘Ÿ๐‘– denote the decomposition level of observational unit ๐‘– and ๐‘ฅ๐‘– its income. The ideology is to model sampling from a finite population as ๐‘–๐‘–๐‘‘ draws from a distribution ๐บ of decomposition levels, ๐‘Ÿ and income, ๐‘ฅ. The decomposition levels could rank the cases with a particular value of ๐‘ฅ. The equivalent income associated with an individual rank of ๐‘Ÿ = 2 may not necessarily count twice as much as incomes associated with an individual rank of ๐‘Ÿ = 1, it may count in fraction. Let ๐ป๐‘› denotes the distribution function of income that results after the decomposition levels have been taken to consideration. Let Mean income = โˆซ ๐‘ฅ ๐‘‘๏ฟฝ๐ป๐‘›(๐‘ฅ)๏ฟฝ๐‘ฅ But ๐‘‘๏ฟฝ๐ป๐‘›(๐‘ฅ)๏ฟฝ = โˆซ ๐‘Ÿ ๐‘‘๏ฟฝ๐บ(๐‘Ÿ,๐‘ฅ)๏ฟฝ๐‘Ÿ โˆซ โˆซ ๐‘Ÿ ๐‘‘๏ฟฝ๐บ(๐‘Ÿ, ๐‘ฅ)๏ฟฝ๐‘Ÿ๐‘ฅ Hence ๏ฟฝ ๐‘ฅ๐‘‘๏ฟฝ๐ป๐‘›(๐‘ฅ)๏ฟฝ = โˆซ โˆซ ๐‘Ÿ๐‘ฅ ๐‘‘๏ฟฝ๐บ(๐‘Ÿ, ๐‘ฅ)๏ฟฝ๐‘Ÿ๐‘ฅ โˆซ โˆซ ๐‘Ÿ ๐‘‘๏ฟฝ๐บ(๐‘Ÿ,๐‘ฅ)๏ฟฝ๐‘Ÿ๐‘ฅ๐‘ฅ [4]. 2.5. Description of the mmoon Bootstrap Method The method ๐‘š๐‘š๐‘œ๐‘œ๐‘› would resample from ๐ป๐‘› distribution which is a decomposed version of ๐น๏ฟฝ๐‘› , the procedure provides and estimates different measures of statistical precision for an estimator ๐œƒ๏ฟฝ. Below is the description of how the method works. Suppose a random sample of size ๐‘› is observed from a completely unspecified probability distribution. ๐‘‹๐‘– = ๐‘ฅ๐‘– , ๐‘‹๐‘– โˆผ ๐น, ๐‘– = 1, โ€ฆ . , ๐‘›. ๐‘ฅ๐‘– โˆผ ๐‘–๐‘–๐‘‘ 1. Construct the sample probability distribution ๐น๏ฟฝ, putting mass 1 ๐‘›๏ฟฝ at each point ๐‘ฅ1 , โ€ฆ ๐‘ฅ๐‘› . ๐น๏ฟฝ is an empirical distribution function (๐ธ๐ท๐น). 2. Stratify the sample into ๐‘Ÿ strata, based on rank of individual ๐‘ฅโ€™๐‘ . 147 American Scientific Research Journal for Engineering, Technology, and Sciences (ASRJETS) (2015) Volume 14, No 1, pp 142-155 3. Fix ๐ป๐‘›, as described in decomposition of empirical distribution. 4. With ๐ป๐‘› fixed, draw a random sample of size ๐‘› ; ๐‘› โƒ›< n, with replacement from ๐ป๐‘›, proportionally from each stratum with respect to ๐‘›๐‘Ÿ ๐‘›๏ฟฝ . This is the bootstrap sample (say ๐‘‹ โˆ— = ๐‘ฅโˆ—). 5. Approximate the sampling distribution of ๐ป๐‘› by the bootstrapping distribution ๐ปโˆ— = ๐ป๐‘› ๏ฟฝ๐‘‹โˆ—,๐น๏ฟฝ๐‘›๏ฟฝ. 6. Repeated realizations of ๐‘‹โˆ— are generated, producing ๐‘‹๐‘˜โˆ— (= ๐‘ฅ1โˆ—, ๐‘ฅ2โˆ—, ๐‘ฅ3โˆ— , โ€ฆ . .. ๐‘ฅ๐‘›โˆ—), such that ๐‘› < ๐‘›;๐‘› โ†’ โˆž ๐‘Ž๐‘›๐‘‘ ๐‘› ๐‘› โ†’ 0 ๐‘Ž๐‘  ๐‘› โ†’ โˆž . Where ๐‘‹๐‘˜โˆ— = ๐‘‹1โˆ—, ๐‘‹2โˆ—, โ€ฆ . .. ๐‘‹1000โˆ— . (i.e. ๐‘˜ independent bootstrap samples, each consisting of ๐‘› data drawn with replacement). Evaluating ๐‘‹โˆ— will produce ๐œƒ๏ฟฝโˆ—. 7. Having chosen a particular ๐œƒ๏ฟฝ (say the mean), obtain empirical bootstrap distribution of ๐œƒ๏ฟฝโˆ—; (๐œƒ๏ฟฝ1โˆ—, ๐œƒ๏ฟฝ2โˆ—, โ€ฆ . .. ๐œƒ๏ฟฝ1000โˆ— ). Precision of the estimator could be tested by obtaining: (i) Bootstrap estimate of standard errors: After evaluating the corresponding bootstrap replications, estimate the standard error of ๐œƒ๏ฟฝ by the empirical standard deviation of the ๐‘˜ replications, the bootstrap estimate of the standard error denoted by ๐‘ ๐‘’๏ฟฝ ๐œƒ๏ฟฝโˆ— is ๐‘ ๐‘’๏ฟฝ ๐œƒ๏ฟฝโˆ— = ๏ฟฝ โˆ‘ ๏ฟฝ๐œƒ๏ฟฝ๐‘˜โˆ— โˆ’ ๏ฟฝฬ…๏ฟฝโˆ—๏ฟฝ2๐พ ๐‘˜=1 ๐พ โˆ’ 1 ๏ฟฝ ๏ฟฝ 1/2 [1] where: ๏ฟฝฬ…๏ฟฝโˆ— = โˆ‘ ๐œƒ๏ฟฝ๐‘˜โˆ—๐พ ๐‘˜=1 ๐พ๏ฟฝ The limit of ๐‘ ๐‘’๏ฟฝ ๐œƒ๏ฟฝโˆ— as ๐พ goes to infinity is the ideal bootstrap estimate of ๐‘ ๐‘’ ๐œƒ : lim๐‘˜โ†’โˆž ๐‘ ๐‘’๏ฟฝ ๐œƒ๏ฟฝโˆ— = ๐‘ ๐‘’ ๐œƒ (ii) Bootstrap estimate of Coefficient Variation: The coefficient of variation of a random variable is defined to be the ratio of its standard error to the absolute value of its mean. The bootstrap coefficient of variation denoted by CV (๐œƒ๏ฟฝโˆ—) refers to the variation at the resampling (bootstrap) level and at population sampling level. CV (๐œƒ๏ฟฝโˆ—) = ๐‘ ๐‘’๏ฟฝ ๐œƒ๏ฟฝโˆ— ๏ฟฝฬ…๏ฟฝโˆ—๏ฟฝ (iii) Bootstrap estimate of bias; Bias is the difference between the expectation of an estimator ๐œƒ๏ฟฝ and the quantity ๐œƒ being estimated. The bootstrap estimate of bias based on the ๐‘˜ replications is: ๐ต๐šค๐‘Ž๐‘ ๏ฟฝ = โˆ‘ ๏ฟฝ๐œƒ๏ฟฝ๐‘˜โˆ— โˆ’ ๏ฟฝฬ…๏ฟฝโˆ—๏ฟฝ๐พ ๐‘˜=1 ๐พ๏ฟฝ 148 American Scientific Research Journal for Engineering, Technology, and Sciences (ASRJETS) (2015) Volume 14, No 1, pp 142-155 where: ๏ฟฝฬ…๏ฟฝโˆ— = โˆ‘ ๐œƒ๏ฟฝ๐‘˜โˆ—๐พ ๐‘˜=1 ๐พ๏ฟฝ (iv) The RMSE of an estimator ๐œƒ๏ฟฝ for ๐œƒ, is ๏ฟฝ๐ธ ๏ฟฝ๏ฟฝ๐œƒ๏ฟฝ โˆ’ ๐œƒ๏ฟฝ 2๏ฟฝ = ๏ฟฝ๐‘ ๐‘’๏ฟฝ๐œƒ๏ฟฝ๏ฟฝ 2 + ๐‘๐‘–๐‘Ž๐‘ ๏ฟฝ๐œƒ๏ฟฝ,๐œƒ๏ฟฝ 2 = ๐‘ ๐‘’๏ฟฝ๐œƒ๏ฟฝ๏ฟฝ.๏ฟฝ1 + ๏ฟฝ๐‘๐‘–๐‘Ž๐‘  ๐‘ ๐‘’ ๏ฟฝ 2 when ๐‘๐‘–๐‘Ž๐‘  = 0, then RMSE = SE (minimum value) [14]. 2.6. Simulation Study This study makes use of simulated data sets drawn from the Singh-Maddala distribution, which can quite successfully mimic observed income distributions in various countries. Two sets of simulation were done for large sample and moderate sample such that ๐‘› is 15000 and 500 respectively. The simulation mimic the parameter values in [3], such that ๐‘Ž = 100, ๐‘ = 2.8, ๐‘ = 1.7 (where a, b, and c are defined above). The values of ๐‘š and ๐‘› are chosen as 20% of original ๐‘› as suggested in [8] and the values increased asymptotically. Authors in [15] suggested choosing replication large enough to minimize statistical error, however statistical error is unavoidable in most situations. Therefore, the number of bootstrap replications chosen in this study is ๐‘˜ = 1000. 3. Results The results are presented in tables 1 & 2 and figures 1 to 8. Figure 1: Chart of Standard Error in Large Sample 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 0 5000 10000 15000 20000 25000 30000 n Standard error,l moon mmoon 149 American Scientific Research Journal for Engineering, Technology, and Sciences (ASRJETS) (2015) Volume 14, No 1, pp 142-155 Table 1: Summary of Bootstrap Estimates in Large Sample n Standard Error Coefficient of variation RMSE Absolute bias Moon mmoon moon mmoon moon mmoon moon mmoon 3000 0.86922 0.22806 0.000316 8.29E-05 0.027487 0.002339 9.08E-14 1.70E-13 4500 0.746153 0.186148 0.000271 6.76E-05 0.023595 0.001909 3.45E-14 8.80E-14 6000 0.633219 0.156741 0.00023 5.69E-05 0.020024 0.001622 1.37E-13 2.89E-14 7500 0.594175 0.145941 0.000216 5.30E-05 0.018789 0.001441 8.80E-14 8.05E-14 9000 0.535212 0.1363 0.000194 4.95E-05 0.016925 0.001303 2.14E-14 5.71E-14 10500 0.493749 0.126693 0.000179 4.60E-05 0.015614 0.001244 5.60E-15 1.14E-13 12000 0.452659 0.115947 0.000164 4.21E-05 0.014314 0.001169 7.60E-14 1.22E-13 13500 0.43135 0.112822 0.000157 4.10E-05 0.01364 0.001063 3.77E-14 1.07E-13 15000 0.405829 0.105592 0.000147 3.84E-05 0.012833 0.001061 5.22E-14 1.37E-13 16500 0.408741 0.099758 0.000148 3.62E-05 0.012926 0.000995 4.41E-14 2.26E-14 18000 0.384636 0.0927 0.00014 3.37E-05 0.012163 0.000961 4.36E-14 2.28E-14 19500 0.350265 0.09395 0.000127 3.41E-05 0.011076 0.000936 7.00E-14 4.85E-14 21000 0.650492 0.0887 0.000127 3.22E-05 0.02057 0.000891 2.05E-13 1.39E-13 22500 0.678618 0.085311 9.02E-05 3.10E-05 0.02146 0.000856 6.29E-14 9.66E-14 24000 0.711078 0.082506 7.25E-05 3.00E-05 0.022486 0.000823 8.01E-13 4.80E-14 Figure 2: Chart of Coefficient of Variation in Large Sample 0 0.00005 0.0001 0.00015 0.0002 0.00025 0.0003 0.00035 0 10000 20000 30000 n Coefficient of Variation,l moon mmoon 150 American Scientific Research Journal for Engineering, Technology, and Sciences (ASRJETS) (2015) Volume 14, No 1, pp 142-155 Figure 3: Chart of RMSE in Large Sample Figure 4: Chart of Absolute Bias in Large Sample Figure 5: Chart of Standard Error in Moderate Sample 0 0.005 0.01 0.015 0.02 0.025 0.03 0 10000 20000 30000 n RMSE,l moon mmoon 0.00E+00 1.00E-13 2.00E-13 3.00E-13 4.00E-13 5.00E-13 6.00E-13 7.00E-13 8.00E-13 9.00E-13 0 10000 20000 30000 n Absolute Bias,l moon mmoon 0 1 2 3 4 5 0 200 400 600 800 1000 n Standard error,m moon mmoon 151 American Scientific Research Journal for Engineering, Technology, and Sciences (ASRJETS) (2015) Volume 14, No 1, pp 142-155 Table 2: Summary of Bootstrap Estimates in Moderate Sample n Standard Error Coefficient of variation RMSE Absolute bias moon mmoon moon mmoon moon mmoon moon mmoon 100 4.455687 1.16504 0.001661 0.000435 0.140901 0.036842 1.66E-13 2.26E-14 150 3.635785 0.930918 0.001354 0.000345 0.114974 0.029438 4.80E-14 3.50E-14 200 3.15536 0.798848 0.001178 0.000299 0.099781 0.025262 8.20E-14 2.87E-14 250 2.764303 0.716033 0.001031 0.000268 0.087415 0.022643 1.45E-13 1.89E-14 300 2.40943 0.656948 0.0009 0.000246 0.076193 0.020775 1.10E-13 8.71E-15 350 2.414765 0.599584 0.0009 0.000224 0.076362 0.018961 2.04E-14 8.46E-15 400 2.170816 0.572092 0.000809 0.000213 0.068647 0.018091 1.24E-13 2.19E-14 450 2.193489 0.526211 0.000818 0.000196 0.069364 0.01664 2.65E-14 8.78E-15 500 1.914705 0.514097 0.000714 0.000191 0.060548 0.016257 6.67E-14 1.48E-13 550 1.838625 0.474448 0.000685 0.000177 0.058142 0.015003 6.01E-15 5.10E-14 600 1.783116 0.462876 0.000665 0.000173 0.056387 0.014637 2.16E-14 3.51E-14 650 1.731724 0.441068 0.000658 0.000164 0.054762 0.013948 7.00E-14 1.36E-13 700 3.463449 0.420461 0.000127 3.22E-05 0.02057 0.000891 2.05E-13 1.39E-13 750 3.018738 0.427549 9.02E-05 3.10E-05 0.02146 0.000856 6.29E-14 9.66E-14 800 2.864389 0.405739 7.25E-05 3.00E-05 0.022486 0.000823 8.01E-13 4.80E-14 Figure 6: Chart of Coefficient of Variation in Moderate Sample 4. Discussion Tables 1 and 2 show asymptotic test results for large and moderate samples respectively. Their corresponding graphs are presented in Figure 1to 8. Both the tables and the figures show some statistical measures of precision on the two bootstrap methods. 0 0.0005 0.001 0.0015 0.002 0 200 400 600 800 1000 n Coefficient of Variation,m moon mmoon 152 American Scientific Research Journal for Engineering, Technology, and Sciences (ASRJETS) (2015) Volume 14, No 1, pp 142-155 From tables 1 and 2, it is revealed that ๐‘š๐‘š๐‘œ๐‘œ๐‘› bootstrap method performs better than ๐‘š๐‘œ๐‘œ๐‘› bootstrap method in large and moderate samples respectively. The ๐‘š๐‘š๐‘œ๐‘œ๐‘› bootstrap method produces smaller estimates of standard error (SE); root mean square error(RMSE); coefficient of variation(CV);absolute bias(ABSB) when compared to the ๐‘š๐‘œ๐‘œ๐‘› bootstrap method. In large sample, when n = 3000: SE = (0.86922;0.22806); CV = (0.000316;8.28E-05); RMSE = (0.027487;0.002339); ABSB = (9.08E-14;1.70E-13) for moon and mmoon respectively and when n = 1800: SE = (0.384636;0.0927); CV = (0.00014;3.37E-05); RMSE = (0.012163;0.000961); ABSB = (4.36E-14;2.28E-14). Also for moderate sample, when n = 100: SE = (4.455687;1.16504); CV = (0.001661;0.000435); RMSE = (0.140901;0.036842); ABSB = (1.66E-13;2.26E-14) for moon and mmoon respectively and when n = 300: SE = (2.40943;0.656948); CV = (0.009;0.000246); RMSE = (0.076193;0.020775); ABSB = (1.10E-13;8.71E-15). Figure 7: Chart of RMSE in Moderate Sample Figure 8: Chart of Absolute Bias in Large Sample The absolute bias of ๐‘š๐‘š๐‘œ๐‘œ๐‘› bootstrap method in large sample is lesser compared to moderate sample, so also the estimate of other measures, these confirmed that, increasing number of samples can reduce effects of random sampling errors which can also arise from bootstrap procedure itself., hence ๐‘š๐‘š๐‘œ๐‘œ๐‘› is suggested as a preferred method in practice. 0 0.02 0.04 0.06 0.08 0.1 0.12 0.14 0.16 0 200 400 600 800 1000 n RMSE,m moon mmoon 0.00E+00 5.00E-14 1.00E-13 1.50E-13 2.00E-13 0 200 400 600 800 1000 n Absolute bias,m moon mmoon 153 American Scientific Research Journal for Engineering, Technology, and Sciences (ASRJETS) (2015) Volume 14, No 1, pp 142-155 The charts in figures 1 to 8 show the bootstrap estimate of SE, CV, RMSE and ABSB. The ๐‘š๐‘š๐‘œ๐‘œ๐‘› bootstrap method converges at a faster rate and it is asymptotically consistent. The ๐‘š๐‘œ๐‘œ๐‘› bootstrap method converges at a slower rate, until ๐‘š becomes as large as the original ๐‘› (i.e. m = n = 15000. See fig. 1 to 8). But when ๐‘š becomes larger than ๐‘›, ๐‘š๐‘œ๐‘œ๐‘› diverges, while ๐‘š๐‘š๐‘œ๐‘œ๐‘› still converges even after ๐‘› becomes larger than ๐‘›. Hence, ๐‘š๐‘š๐‘œ๐‘œ๐‘› is asymptotically consistent. 5. Conclusion The ๐‘š๐‘š๐‘œ๐‘œ๐‘› bootstrap method for heavy-tailed distributions in large and moderate samples is justified through its validity and application to empirical work. The bootstrap estimates of standard error and other measures of statistical precision (such as absolute bias, coefficient of variation and root mean squared error) confirmed the reliability and suitability of the method in practice. 6. Recommendation Bootstrap researchers interested in heavy-tailed distributions should decompose the empirical distribution before resampling so as to overcome the difficulties posed by outliers. References [1] M. R. Chernicks. Bootstrap Methods: A guide for Practitioners and Researchers. Hoboken, New Jersey: John Wiley & Sons, Inc., 2008, pp. 1, 50, 180. [2] D. Brownstone and R. Valletta. โ€œThe Bootstrap and Multiple Imputations.โ€ Internet: http://www.economics.uci.edu/~dbrownst/bootmi.pdf. Dec. 28, 2000. [Feb. 07, 2012]. [3] R. Davidson and E. Flachaire. (2007, Oct.). โ€œAsymptotic and Bootstrap Inference for Inequality and Poverty Measures.โ€ Journal of Econometrics. [ On-line]. 141(1), pp. 141-166. Available: https://hal.archives- ouvertes.fr/halshs-00175929/document [Mar. 19, 2012]. [4] M. Biewen (2001, Nov.). โ€œBootstrap Inference for Inequality, Mobility and Poverty Measurement.โ€ Journal of Econometrics. [On-line]. 108 (2002), pp. 317-342. Available: http://kumlai.free.fr/RESEARCH/THESE/TEXTE/MOBILITY/mobility%20salariale/Bootstrap%20inference% 20for%20inequality,%20mobility%20and%20poverty%20measurment.pdf [Dec. 07, 2011]. [5] J. A. Mills and S. Zandvakili. โ€œStatistical Inference via Bootstrapping for Measures of inequalityโ€ Internet: http://www.levyinstitute.org/publications/statistical-inference-via-bootstrapping-for-measures-of-inequality, Feb. 1997. [Jun. 06, 2012] [6]Wikimedia Foundation, Inc. โ€œBootstrapping (Statistics).โ€ Internet: http://en.wikipedia.org/wiki/Bootstrapping_(Statistics), Aug. 18, 2015 [Oct. 28, 2013]. 154 https://hal.archives-ouvertes.fr/halshs-00175929/document https://hal.archives-ouvertes.fr/halshs-00175929/document http://kumlai.free.fr/RESEARCH/THESE/TEXTE/MOBILITY/mobility%20salariale/Bootstrap%20inference%20for%20inequality,%20mobility%20and%20poverty%20measurment.pdf http://kumlai.free.fr/RESEARCH/THESE/TEXTE/MOBILITY/mobility%20salariale/Bootstrap%20inference%20for%20inequality,%20mobility%20and%20poverty%20measurment.pdf http://www.levyinstitute.org/publications/statistical-inference-via-bootstrapping-for-measures-of-inequality,%20Feb.%201997 http://www.levyinstitute.org/publications/statistical-inference-via-bootstrapping-for-measures-of-inequality,%20Feb.%201997 http://en.wikipedia.org/wiki/Bootstrapping_(Statistics) American Scientific Research Journal for Engineering, Technology, and Sciences (ASRJETS) (2015) Volume 14, No 1, pp 142-155 [7] A. Antoniadis. โ€œBootstrap Methods: Recent Advances and New Applications.โ€ Internet: http://www. mescal.imag.fr/membres/yves.denneulin/LASCAR/page2/files/BootstrapAA.pdf, Oct. 2007. [Jun. 14, 2012] [8] K. Singh and M. Xie. โ€œBootstrap: A Statistical Method.โ€ Internet: www.stat.rutgers.edu/home/mxie/rcpapers /bootstrap.pdf, [Feb. 08, 2012]. [9] P. J. Bickel and A. Sakov (2002, Apr.). โ€œExtrapolation and the boostrap.โ€ The Indian Journal of Statistics. [On-line]. 64(3), pp. 640- 652. Available: http://www.jstor.org/stable/25051419 [Nov. 13, 2014]. [10] S. M. S. Lee and M. C. Pun (2006, Sep.). โ€œOn ๏ฟฝโˆ’๏ฟฝ๏ฟฝ๏ฟฝ ๏ฟฝ๏ฟฝโˆ’๏ฟฝ Boostrapping for Non Standard M- estimation with Nuisance Parameters.โ€ Journal of American Statistics Association. [On-line]. 101(475), pp. 1185-1197. Available: http://www.jstor.org/stable/27590794 [Jan. 20, 2015]. [11] P. J. Bickel and D. A. Freedman (1981, Mar.). โ€œSome Asymptotic Theory for the Bootstrap.โ€ The Annals of Statistics. [On-line]. 9(6), pp. 1196-1217. Available: http://projecteuclid.org/download/pdf_1/euclid.aos/1176345637, [Feb. 08, 2012]. [12] B. Sen, M. Banerjee and M. Woodroofe (2010, Oct.). โ€œInconsistency of Bootstrap: The Grenander Estimator.โ€ The Annals of Statistics. [On-line]. 38(4), pp. 1953-1977. Available: http://arxiv.org/abs/1010.3825, [Feb. 17, 2012]. [13] M. R Chernicks and R. A. LaBudde. (2011, Nov. 1). An Introduction to Bootstrap Methods with Applications to R. (1st edition) [On-line]. Available: http://nnpdf.pillaroftheworld.com/an-introduction-to- bootstrap-methods-robert-a-labudde-96621231.pdf [Jun. 14, 2012]. [14] T.O Olatayo, G.N. Amahia and T.O. Obilade. ( 2010). โ€œBootstrap Method for Dependent Data Structure and Measure of Statistical Precision.โ€ .Journal of Mathematics and Statistics. [On-line]. 6(2), pp. 84-88. Available: http://thescipubl.com/abstract/10.3844/jmssp.2010.84.88, [Jun. 08, 2012]. [15] A. C. Davison and D. Kuonen. (2002). โ€œAn Introduction to the Bootstrap with Applications in R.โ€ Statistical Computing and Graphics Newsletter. [On-line]. 13(1), pp. 6-11. Available: http://www.statoo.com/en/publications/bootstrap_scgn_v131.pdf [Aug. 22, 2012]. 155 http://www.stat.rutgers.edu/ http://www.jstor.org/stable/27590794 http://arxiv.org/abs/1010.3825 http://thescipubl.com/abstract/10.3844/jmssp.2010.84.88 = ๐‘ ๐‘’,,๐œƒ...,1+,,,๐‘๐‘–๐‘Ž๐‘ -๐‘ ๐‘’..-2..