compiles/8f0013ef0f2b9c6f92df168d9a909d7c/output.dvi EUROPEAN JOURNAL OF PURE AND APPLIED MATHEMATICS Vol. 7, No. 1, 2014, 22-36 ISSN 1307-5543 – www.ejpam.com Empirical Likelihood Ratio Based Goodness-of-Fit Test for the Generalized Lambda Distribution Wei Ning Department of Mathematics and Statistics, Bowling Green State University, Bowling Green, OH, USA Abstract. In this paper, we propose a goodness-of-fit test based on the empirical likelihood method for the generalized lambda distribution (GLD) family. Such a nonparametric test approximates the optimal Neyman-Pearson likelihood ratio test under the unknown alternative distribution scenario. The p-value of the test is approximated through the simulations due to the dependency of the test statistic on the data. The test is applied to the roller data set and the pollen data set to illustrate the testing procedure for the sufficiency of the GLD fittings. 2010 Mathematics Subject Classifications: 62G05, 62G10, 62G20 Key Words and Phrases: Generalized lambda distribution; Empirical likelihood; Nonparametric; Goodness-of-fit test. 1. Introduction Modeling the skewed or the heavy tailed data is an important issue in statistical data fitting. There are extensive distribution families proposed by many researchers to achieve this goal. The generalized lambda distribution (GLD) family was originally introduced by Tukey [20], who proposed an one-parameter lambda distribution. Tukey’s lambda distribution was generalized, for the purpose of generating random variables for Monte Carlo simulation studies, to the four parameters GLD proposed by Ramberg and Schmeiser [13, 14]. Ramberg et al. [15] developed a four-parameter system with the tables for fitting a wide variety of curve shapes. Since the early 1970s, the GLD has been applied to fitting phenomena in many fields of endeavor with continuous probability density functions (pdf). The GLD family with four parameters λ1, λ2, λ3, λ4, which is denoted as GLD(λ1,λ2,λ3,λ4), has a probability density function f (x) = λ2 λ3 yλ3−1+λ4(1− y)λ4−1 , at x =Q(y), (1) Email address: wning@bgsu.edu http://www.ejpam.com 22 c© 2014 EJPAM All rights reserved. W. Ning / Eur. J. Pure Appl. Math, 7 (2014), 22-36 23 where Q(y) is the percentile function defined as Q(y) = λ1+ yλ3 − (1− y)λ4 λ2 , where 0 ≤ y ≤ 1, λ1 and λ2 are location and scale parameters respectively. Karian and Dudewicz [4] gave the first four moments of GLD(λ1,λ2,λ3,λ4) with λ3 >−1/4, λ4 > −1/4 as α1 = µ= E(X ) = λ1+ A λ2 , α2 = σ 2 = E[(X −µ)2] = B− A2 λ2 2 , α3 = E(X − E(X ))3/σ3 = C − 3AB+ 2A3 λ3 3σ 3 , α4 = E(X − E(X ))4/σ4 = D− 4AC + 6A2B− 3A4 λ4 2σ 4 (2) where A= 1 1+λ3 − 1 1+λ4 , B = 1 1+ 2λ3 + 1 1+ 2λ4 − 2β(1+λ3, 1+λ4), C = 1 1+ 3λ3 − 1 1+ 3λ4 − 3β(1+ 2λ3, 1+λ4) + 3β(1+λ3, 1+ 2λ4), D = 1 1+ 4λ3 + 1 1+ 4λ4 − 4β(1+ 3λ3, 1+λ4) + 6β(1+ 2λ3, 1+ 2λ4) − 4β(1+λ3, 1+ 3λ4), and β(·, ·) is a Beta function. The GLD family is known for its high flexibility on approaching many well-known distributions and ability to fit the data sets with different shapes, especially with those with heavy tails. There has been further extensive work done in this field. For example, Karian and Dudewicz [4] provided the tabulated tables for the method of moment and the percentile method which are used to estimate the parameters of the GLD. King and MacGillivray [6] and Lakhany and Massuer [7] considered a definite fit to the data set by maximizing the goodness of fit. Su [16] used a discretized method to fit GLD to the empirical data. Su [17] derived the estimation procedure by using the maximum likelihood method. Asquith [1] provided L-moments and TL-moments for the GLD. Fournier et al. [2] proposed a new estimation method by combining the method of moment and the percentile method. Ning et al. [8] considered the fitting problem involving the mixture of two GLDs and as a result made the comparisons to the other mixture distribution families. Ning and Gupta [9] proposed a GLD change point model to detect the change points for the DNA copy number. Su et al. [19] proposed a GLD based calibration model to achieve more flexible data fitting comparing the classic normal calibration model and the skew normal calibration model. For W. Ning / Eur. J. Pure Appl. Math, 7 (2014), 22-36 24 the other recent work related to the GLD family and its applications, the readers are referred to Karian and Dudewicz [5]. As for data fitting, the extent to which how well the proposed distribution family can fit the data is an important issue. Goodness-of-fit test is the test that is always used to check the sufficiency of the data fitting by a given distribution. Neyman-Pearson lemma indicates that the likelihood ratio test (LRT) is the uniformly powerful (UMP) test for the hypotheses: H0 : f = f0 versus H1 : f = f1 with f0 and f1 both known. However, the alternative distribu- tion is usually not known in practice. Recently, Vexler and Gurevich [22] and Vexler et al. [23] constructed goodness-of-fit test based on the empirical likelihood method to approximate the optimal Neyman-Pearson likelihood ratio test with an unknown alternative density function. In this paper, we will consider the goodness-of-fit test for the generalized lambda distri- bution (GLD) family. We will adopt the idea as that of Vexler and Gurevich [22] similarly to construct a nonparametric goodness-of-fit test to test the null hypothesis of a GLD versus the alternative hypothesis of some other unknown distribution. This paper is organized as fol- lows. In Section 2, a brief introduction of the empirical likelihood method will be given. The empirical likelihood ratio based goodness-of-fit test is proposed and its asymptotic properties will be derived. In Section 3, an empirical procedure based on the simulations is provided to approximate the p-value of the test statistic in Section 2 due to the dependency of the test statistic on the estimated parameters. The proposed goodness-of-fit test is applied to the roller data set and the pollen data set in Section 4 to illustrate the testing procedure and fitting results are given. Discussion is provided in Section 5. 2. Statistical Method 2.1. The Empirical Likelihood (EL) Method Consider the independently and identically distributed p-dimensional observations, say x1, · · · , xn, from an unknown population distribution F . The main idea of empirical likelihood methods, proposed and systematically developed by Owen [10] is to place a probability mass at each observation. Therefore, let pi = P(X = x i) and the empirical likelihood function of F be defined as L(F) = n∏ i=1 pi . It is clear that L(F) subject to the constraints pi ≥ 0 and ∑ i pi = 1 is maximized at pi = 1/n, i.e., the likelihood L(F) attains its maximum n−n under the full nonparametric model. When a population parameter θ identified by Em(X ,θ ) = 0 is of interest where m(x ,θ ) is a real-valued function, the empirical log-likelihood maximum when θ has the true value θ0 is obtained subject to the additional constraint ∑ pim(x i ,θ0) = 0. W. Ning / Eur. J. Pure Appl. Math, 7 (2014), 22-36 25 The empirical log-likelihood ratio statistic to test θ = θ0 is given by R(θ0) =max{ ∑ i log npi : pi ≥ 0, ∑ pi = 1, ∑ pim(x i ,θ0) = 0}. Owen [10, 11] shows that similarly to the likelihood ratio test statistic in a parametric model setup, with mild regular conditions θ0, −2 log R(θ0)→ χ 2 r in distribution under the null model θ = θ0, where r is the dimension of m(x ,θ ). More details of the empirical likelihood and the related work refer to Owen [12]. 2.2. The EL Goodness-of-Fit Test We will test the following hypothesis: H0 : f = f0 ∼ GLD(λ1,λ2,λ3,λ4) H1 : f = f1 � GLD(λ1,λ2,λ3,λ4). The likelihood ratio test statistic for this hypothesis is defined as LR= ∏n i=1 fH1 (x i)∏n i=1 fH0 (x i) = ∏n i=1 fH1 (x i)∏n i=1 f (x i |λ) where x1, x2, · · · , xn follows a GLD distribution with the parameter λ= (λ1,λ2,λ3,λ4) under the null hypothesis. Neyman-Pearson lemma guarantees that such a test is the UMP test with f0 and f1 both known. If they are both unknown, the maximum likelihood method will be applied to estimate the parameters λ̂ = (λ̂1, λ̂2, λ̂3, λ̂4) of a GLD distribution under the null hypothesis. The parameters can be estimated by using the R package GLDEX developed by Su [18]. We then will apply maximum empirical likelihood method to estimate the numerator. We rewrite L f = n∏ i=1 fH1 (x i) = n∏ i=1 fH1 (x(i)) = n∏ i=1 fi , where x(1) ≤ x(2) ≤ · · · ≤ x(n) are the order statistics of the observations x1, · · · , xn. We will apply the empirical likelihood method introduced in Section 2.1 to derive the values of fi to maximize L f with the constraint ∫ f (s)ds = 1 corresponding to the alternative hypothesis. We first give the following lemma by Vexler and Gurevich [22] to express this constraint more explicitly. Lemma 1. Let X1, · · · , Xn be independent and identically distributed random variables with a density function f (x). Then n∑ j=1 ∫ X( j+m) X( j−m) f (x)d x =2m ∫ X(n) X(1) f (x)d x − m−1∑ k=1 (m− k) ∫ X(n−k+1) X(n−k) f (x)d x − m−1∑ k=1 (m− k) ∫ X(k+1) X(k) f (x)d x t 2m ∫ X(n) X(1) f (x)d x − m(m− 1) n , (3) W. Ning / Eur. J. Pure Appl. Math, 7 (2014), 22-36 26 where X( j) = X(1) if j ≤ 1, and X( j) = X(n), if j ≥ n. X(1) < X(2) < · · · < X(n) are order statistics of X1, · · · , Xn. See Vexler and Gurevich [22] for the detailed proof. Since ∫ X(n) X(1) f (x)d x ≤ ∫∞ −∞ f (x)d x = 1, from Lemma 1 we have Λm ≤ 1,Λm = 1 2m n∑ j=1 ∫ X( j+m) X( j−m) f (x)d x (4) Therefore, Λm → 1 as m n → ∞. The integration on the right side of the equation (4) can be approximated as, n∑ j=1 ∫ X( j+m) X( j−m) f (x)d x t (X( j+m)− X( j−m)) f (x( j)) = (X( j+m)− X( j−m)) f j . Thus, Λm t 1 2m n∑ j=1 (X( j+m)− X( j−m)) f j ¬ eΛm, therefore, eΛm ≤ 1. To maximize ∏ f j with this constraint, we apply the Lagrange multiplier method and have l( f1, · · · , fn,η) = n∑ j=1 log f j +η( 1 2m n∑ j=1 (X( j+m)− X( j−m)) f j − 1), (5) where η is a lagrange multiplier. By taking the derivative of the equation (5) respect to each f j and η, we obtain ∂ l ∂ fi =0⇒ 1 f j + η 2m (X( j+m)− X( j−m)) = 0 (6) ∂ l ∂ η =0⇒ 1 2m n∑ j=1 (X( j+m)− X( j−m)) f j − 1= 0. (7) from the equation (6) and (7), we have, ∑ f j · 1 f j +η 1 2m ∑ f j(X( j+m)− X( j−m)) = 0⇒ η= −n. (8) Hence, we will obtain the estimate of f j to maximize ∏ f j as f j = 2m n(X( j+m)− X( j−m)) , (9) W. Ning / Eur. J. Pure Appl. Math, 7 (2014), 22-36 27 where X( j) = X(1), if j ≤ 1, and X( j) = X(n), if j ≥ n. We then construct the likelihood ratio test statistic for the goodness-of-fit test for the GLD based on the maximum empirical likelihood method as GLDmn = ∏n j=1 2m n(X( j+m)−X( j−m)) max λ ∏n j=1 fH0 (X j|λ) , (10) where λ = (λ1,λ2,λ3,λ4) is the parameter vector of a GLD. We notice that the test statistic GLDmn strongly depends on the integer m. To make the test more efficient, Vexler and Gure- vich [22] and Vexler et al. [23] reconstructed the test statistic according to the properties of the empirical likelihood method. We follow their suggestions here to reconstruct the test statistic in (10) as GLDn = min 1≤m 0 as n→∞. (A11)