Academic Journal of Science and Technology ISSN: 2771-3032 | Vol. 5, No. 1, 2023 112 Research on Two‐stage Estimation of Partially Linear Single‐index Model with Longitudinal Data Chaojie Chang School of Mathematics and Science, China University of Geosciences (Wuhan), Wuhan, China Abstract: Partial linear single-index model is a kind of semi-parametric model with wide application. In this paper, we deal with the partial linear single-index model under longitudinal data. A "two-stage estimation method" without iteration by using local polynomial and bias correction generalized estimation equation is proposed. under some regularity conditions, the asymptotic properties of the connection function and unknown parameter estimator are investigated. Numerical simulation shows that the proposed method is robust. Keywords: Partial linear single-index model, Longitudinal data, Local polynomial, Generalized estimation equation for correcting deviation, Asymptotic normality. 1. Introduction This paper studies the partial linear single-indicator model under the following longitudinal data: π‘Œ 𝑍 πœƒ g 𝑋 𝛽 πœ€ ,𝑖 1, β‹― ,𝑛,𝑗 1, β‹― ,π‘š (1) Among them, the explanatory variables 𝑋 and 𝑍 are related to the model error πœ€ Independent from each other. g βˆ™ is a known nonparametric connection function, 𝛽 is an unknown q-dimension index parameter, πœƒ is an unknown p- dimensional linear parameter, πœ€ is the mean value is 0, and the variance is 0 𝜎 ∞ random error. To ensure the identifiability of the single indicator part, suppose ‖𝛽‖ 1, and the first non-zero element of the parameter vector is positive. In this paper, under the longitudinal data, each individual is observed for a limited number of times, namely π‘š 𝑐, then the total number of observations is βˆ‘ π‘š Where π‘š is a bounded positive integer sequence, and it can be assumed that 𝑛 is infinite, which means that the total number of observations 𝑁 and the total number of subjects n are of the same order, that is, 𝑁 β†’ ∞ and 𝑛 β†’ ∞ are equivalent. Partial linear single-indicator model is the combination of linear model and single-indicator model. Its application range is very wide, which has aroused the research interest of many scholars. Many methods have been proposed to estimate unknown parameters and non-parametric connection functions[1-4]. When π‘ž 1, the model (1) is a partial linear model of longitudinal data [5-6]; When 𝑝 0, the model (1) is a single indicator model of longitudinal data [7-8]; When g βˆ™ selects 0, model (1) is a linear regression model of longitudinal data [9]. The existing estimates of partial linear single-indicator models are mainly based on the mean regression of least squares or likelihood method, and the distribution of random error needs to be assumed. However, if the assumption is incorrect, the estimation results may be inaccurate or even wrong, and the research work will be meaningless. In addition, at present, the linear parameters in the model πœƒ and index parameters 𝛽 is almost estimated at the same time, and requires multiple iterations. The calculation is very complex and time-consuming. Then, due to 𝛽 and πœƒ may have some correlation, which may cause 𝛽 is difficult to identify, which will also lead to inaccurate estimation results. In view of the above problems, without considering the intra-group correlation, this paper proposes a "two-stage estimation method" without iteration based on local polynomial estimation and bias correction generalized estimation equation 𝛽 And linear parameters πœƒ This method can reduce the running time and improve the operation efficiency. In addition, in order to solve 𝛽 and πœƒ For the linear correlation between 𝑋 and 𝑍, the formula (2) is introduced to eliminate the correlation between X and Z, so as to improve the accuracy of parameter estimation. In section 1, we give a "two-stage estimation method" without iteration based on local polynomial estimation and bias correction generalized estimation equation [10]. In section 2, we study the index parameters under some regularity assumptions Ξ² And linear parameters ΞΈ The asymptotic normality of the estimator and the asymptotic normality of the estimator of the connection function g βˆ™ . The random simulation experiment in section 3 shows that the estimator obtained by this method is robust. Assume that the observation value is 𝑋 ,π‘Œ ,𝑍 ;𝑖 1, β‹― ,𝑛,𝑗 1, β‹― ,π‘š from the sample of model (1). For the convenience of calculation: π‘Œ 𝑦 , β‹― ,𝑦 , 𝑋 π‘₯ , β‹― ,π‘₯ 𝑍 𝑧 , β‹― ,𝑧 ,πœ€ πœ€ , β‹― ,πœ€ Then the model (1) can be written as the following vector matrix form: π‘Œ 𝑍 πœƒ 𝑔 𝑋 𝛽 πœ€ ,𝑖 1, β‹― ,𝑛 To ensure the identifiability of the model, it is assumed that 𝛽 1,𝛽 , β‹― ,𝛽 . In order not to lose generality, real parameters are assumed 𝛽 is a positive definite matrix. In addition, based on the linear correlation between 𝑋 and 𝑍,let 𝑍 πœ‘ 𝑋 𝛽 πœ‚ (2) here πœ‘ βˆ™ is an unknown function from 𝑅 to 𝑅 , 𝛽 is a 𝑝 𝑑 with standard orthogonal columns matrix, πœ‚ the mean value of is zero and is independent of 𝑋. The dimension 𝑑 is usually much smaller than the dimension 𝑝 of 𝑋 , 113 which is a common dimensionality reduction assumption in the literature. In this chapter, we conduct statistical inference research under the condition of 𝑑 1. Generally, once you get 𝛽 The βˆšπ‘› of is estimated and inserted (1), and the best estimate can be achieved by the method developed for the partial linear model πœƒ , however, 𝛽 and πœƒ may be related, causing 𝛽 is difficult to identify. This is the advantage of introducing model assumption (2), because it allows deleting the 𝑍 part related to 𝑋, so that the residual in (2) πœ‚ Will be independent of 𝑋 . Similarly, it is necessary to add an identifiability condition, i.e. ‖𝛽 β€– 1 and the first component of the parameter is positive. The parameter vector is given below 𝛽 and πœƒ and connection function g βˆ™ The estimation algorithm of "two- stage estimation method": Algorithm for Stage One: 1. First, construct a regression model 𝑍 and 𝑋 is regressed to obtain 𝛽 estimator of 𝛽 2. Using local smoothing estimation to get οΏ½Μ‚οΏ½ 𝑍 πœ‘ 𝑋 𝛽 ,so π‘Œ οΏ½Μ‚οΏ½ πœƒ β„Ž 𝑋 𝛽 𝑋 𝛽 πœ€ , her πœ‚ and 𝑋 are independent of each other. 3. For π‘Œ and οΏ½Μ‚οΏ½ carries out linear regression to obtain πœƒ initial estimate of πœƒ 4. Construct a regression model π‘Œ 𝑍 πœƒ 与 and 𝑋 is regressed to obtain 𝛽 initial estimate of 𝛽 5. Using local polynomial smoothing estimation, the initial feasible estimator of the connection function g βˆ™ and its first derivative g βˆ™ is obtained g and g 𝑔 𝑒 𝑔 𝑒;πœƒ,𝛽 π‘Š 𝑒;𝛽 π‘Œ 𝑍 πœƒ 𝑔 𝑒 𝑔 𝑒;πœƒ,𝛽 π‘Š 𝑒;𝛽 π‘Œ 𝑍 πœƒ Algorithm for Stage Two: 6. Use Step 4 to get the initial estimate 𝛽 , obtained by solving the deviation correction generalized estimation equation πœƒ initial estimate of πœƒ 1 𝑛 π‘Œ 𝑍 πœƒ 𝑔 𝑋 𝛽 𝑍 𝐸 𝑍 𝑋 𝛽 0 7. Use the updated πœƒ initial estimate of πœƒ form a new residual π‘Œ 𝑍 πœƒ ,Then it is obtained by solving the deviation correction generalized estimation equation as follows 𝛽 initial estimate of 𝛽 1 𝑛 π‘Œ 𝑍 πœƒ 𝑔 𝑋 𝛽 𝑔 𝑋 𝛽 𝑋 𝐸 𝑋 𝑋 𝛽 0 8. Use the updated estimates of πœƒ and 𝛽 in steps 6 and 7 to updated the estimate of g and g ,following the steps as described in step 5. The first stage is to use the linear regression method and local linear smoothing estimation to obtain the initial estimates of unknown parameters, connection functions and their derivatives and residuals. In the second stage, the final estimates of unknown parameters and connection functions are obtained by using the deviation correction generalized estimation equation. The algorithm can obtain the asymptotic property of the unknown parameter estimator without iteration, and solve 𝛽 and πœƒ may be a related problem that leads to inaccurate parameter estimation; In addition, the estimators of the connection function g βˆ™ and its derivative g βˆ™ are also obtained. 2. Main Result In order to study the theoretical results of the estimators obtained by the "two-stage estimation method", the following regularity assumptions are first given: C1 Individuals and individuals are independent and equally distributed. C2 (i) The distribution of 𝑋 has a compact support set π’œ; (ii) The bounded positive function 𝑓 𝑑 is 𝑋 𝛽 density function on 𝑇, and 𝛽 The domain of satisfies the Lipschitz condition of order 1.here 𝑇 𝑑 𝑋 𝛽;𝑋 ∈ 𝐴,𝑖 1, β‹― ,𝑛,𝑗 1, β‹― ,π‘š . C3 (i) Join functions g and g has a second order continuous partial derivative, whereg is the function vector g the ith component of 𝑖 , 1 𝑖 π‘ž,g1 𝑑 𝐸 𝑍𝑖𝑗 𝑋𝑖𝑗 𝑇 𝛽 𝑑 ; (ii) g satisfies the first-order Lipschitz condition, where g is the jth component of g , 1 𝑗 𝑝,𝑔 𝑑 𝐸 𝑋 𝑋 𝛽 𝑑 . C4 On 𝑅 , the kernel function 𝐾 βˆ™ is a continuous bounded probability density function that satisfies the Lipschitz condition and satisfies 𝑒 𝐾 𝑒 𝑑𝑒 0 , |𝑒| 𝐾 𝑒 𝑑𝑒 ∞. C5 (i) it has a constant 𝑀 0 , so that for any𝑖,𝑗 satisfies sup ∈ 𝐸 𝑍 𝑋 𝛽 𝑑 𝑀 ∞; (ii) Random error πœ€ independent of covariate 𝑋 and 𝑍 , and there is a constant 𝑀 0 , πœ€ 0,0 π‘£π‘Žπ‘Ÿ πœ€ 𝜎 ∞,𝐸 πœ€ 𝑀 ∞. C6 When 𝑛 β†’ ∞,bandwidth sequence β„Ž and β„Ž satisfie (i) π‘›β„Ž /log 𝑛 β†’ ∞, lim β†’ π‘›β„Ž ∞; (ii) π‘›β„Ž β„Ž /log 𝑛 β†’ ∞,π‘›β„Ž β†’ ∞. C7 Definition π‘Œβˆ— π‘Œπ‘–π‘— βˆ— 𝑍𝑖𝑗 𝑇 πœƒ0οΌŒπ›΄ is a bounded positive definite matrix 𝛴 πΆπ‘œπ‘£ 𝑍 𝐸 𝑍|𝑋 𝛽 𝐡 𝐸 𝑔 𝑋 𝛽 𝑋 𝑋|𝑋 𝛽 𝑋 𝑋|𝑋 𝛽 𝐡 𝐸 𝑓 βˆ—π‘” 𝑋 𝛽 𝑔 𝑋 𝛽 𝑋 𝑋|𝑋 𝛽 𝑋 𝑋|𝑋 𝛽 Condition (C1) is the assumption of independence. See reference [11] for details. Lipschitz condition and standard smoothness condition in condition (C2) and condition (3). For some common regularity and slip assumptions of conditional (C4) kernel functions, to ensure that the theoretical results are well established. The condition (C5) is to satisfy the existence of the second moment, so that the proposed parameter and the single exponential function estimator are consistent and asymptotically normal. Condition (C6) is the bandwidth h used to estimate the connection function g βˆ™ , and the other band width β„Ž is to control the change of g βˆ™ , see reference [12]. The condition (C7) ensures that the variance limit of the parameter estimator exists. Let 𝜌 𝑒 𝐾 𝑒 𝑑𝑒 ,𝜚 𝐾 𝑒 𝑑𝑒 ,𝑙 1,2,3 ,the following are the connection function g βˆ™ and unknown parameter components πœƒ and 𝛽 Results of asymptotic 114 properties of. Theorem 1 Assume that the above regularity and smoothness conditions C1-C4 hold, and if π‘›β„Ž β†’ 0 , the initial estimator𝛽 ,πœƒ satisfied 𝛽 𝛽 𝑂 𝑛 / , πœƒ πœƒ 𝑂 𝑛 / , then when 𝑛 β†’ ∞ βˆšπ‘›β„Ž 𝑔 𝑒;𝛽,πœƒ 𝑔 𝑒 β„Ž 𝑏 𝑒 β†’ 𝑁 0,𝐴 𝑒 here 𝑏 𝑒 𝑔 𝑒 𝜌 ,𝐴 𝑒 βˆ— 𝑔 𝑒 𝑒 Theorem 2 Assume that the above regularity and smoothness conditions C1-C7 hold, and the initial estimator 𝛽 ,𝛽 satisfied 𝛽 𝛽 𝑂 𝑛 / , 𝛽 𝛽 𝑂 𝑛 / , then when 𝑛 β†’ ∞ βˆšπ‘› πœƒ πœƒ β†’ 𝑁 0,𝜎 𝛴 βˆšπ‘› 𝛽 𝛽 β†’ 𝑁 0,𝜎 𝐡 𝐡 𝐡 where 𝐡 is the inverse of 𝐡 ,the definition of 𝛴、𝐡 and 𝐡 see C7. 3. Simulation Study In this section, this paper considers the effectiveness of the "two-stage estimation method" by applying the Monte Carlo method in the case of limited samples. First, generate some random numbers in the following model: π‘Œ 𝑍 πœƒ 𝑔 𝑋 𝛽 πœ€ ,𝑖 1, β‹― ,𝑛,𝑗 1, β‹― ,π‘š where 𝑍 is a covariate, a 0 1 distribution with a parameter of 0.5, 𝑋 𝑋 ,𝑋 ,𝑋 ,𝑋 ,𝑋 𝑋 、𝑋 、𝑋 、𝑋 and 𝑋 are independent of each other and are uniform distribution from the 0,1 interval. During the simulation experiment, errors πœ€ πœ€ ,πœ€ ,πœ€ ,πœ€ ,πœ€ ,𝑖 1,2,3,4,5 obey standard normal distribution. 𝛽 0.5,0,0.5,0.5,-0.5 ,𝛽 0.75,0.5,-0.25,-0.25,0.2 and πœƒ 1 the connection function 𝑔 𝑒 sin ,𝐴 √ . , 𝐢 √ . . In addition, this paper uses the kernel function 𝐾 π‘₯ √ 𝑒π‘₯𝑝 , β„Ž is the window width, which is selected by the following generalized cross validation method and meets the assumption C6: 𝐺𝐢𝑉 β„Ž 1 𝑛 π‘Œ 𝑍 πœƒ 𝑔 𝑋 𝛽 𝑛 π‘‘π‘Ÿ 𝐼 𝑆 For the convenience of comparison, n is selected as 50 "," 150 and 200 respectively, and the simulation results of each case are based on 500 repeated experiments. In this paper, the following evaluation indicators are used to evaluate the accuracy of parameter estimation: linear parameter ΞΈ Estimator ΞΈ Μ‚ Deviation, standard deviation, mean square error and index parameters of Ξ² Estimator Ξ² Μ‚ The deviation, standard deviation and mean square error of, and the estimator of the linking function g (βˆ™) g Μ‚ (βˆ™) The mean and standard deviation of RMSE, where the estimated RMSE of the connection function g (βˆ™) is calculated by the square root of the following mean square error: RMSE 𝑔 1 π‘šπ‘› 𝑔 𝑋 𝛽 𝑔 𝑋 𝛽 Table 1 below shows the deviation, standard deviation and mean square error of the estimator when n is taken as 50 "," 150 "and 200 respectively. Figure 1 shows the scatter diagram of the connection function when n=50", "150" and 200 ". Table 1. 𝛽 and πœƒ deviation, standard deviation and mean square error of estimator, g (Μ‚βˆ™) Mean and standard deviation of RMSE 𝒏 150 200 300 bias std ste mse bias std ste mse bias std ste mse ΞΈ 0.0035 0.1274 0.1278 0.0163 0.0024 0.1026 0.1131 0.0125 0.0017 0.0894 0.0921 0.008 Ξ²1 0.0432 0.2644 0.251 0.0718 0.0466 0.2372 0.2163 0.0584 0.0311 0.1918 0.1746 0.0377 Ξ²2 0.0209 0.1304 0.1217 0.0174 0.0193 0.1102 0.1053 0.0125 0.0125 0.0898 0.0854 0.0083 Ξ²3 0.0414 0.2144 0.1996 0.0477 0.0413 0.1956 0.1764 0.0506 0.0264 0.1611 0.1443 0.0267 Ξ²4 0.0221 0.2479 0.2253 0.0143 0.0109 0.2196 0.1945 0.0105 0.0311 0.1761 0.0763 0.0068 Ξ²5 0.0418 0.2005 0.1793 0.0419 0.0432 0.1002 0.0939 0.0352 0.0138 0.0815 0.1302 0.0225 me se me se me se g 0.0412 0.0142 0.0333 0.0121 0.0296 0.0076 Figure 1. Scatter diagram of real connection function and estimated connection function when n=50,100,200 From Table 1 and Figure 1, the following conclusions can be drawn: The deviation of parameter estimation, standard error, mean and mean square error of standard error, as well as the mean and standard error of RASE estimated by the connection function all decreased significantly with the 115 increase of samples. Therefore, the "two-stage estimation method" proposed in this paper is relatively stable in the estimation of parameters, and the connection function estimation and the fitting effect of the real curve are good. 4. Conclusion This paper presents a new "two-stage estimation method", which is based on local polynomial and bias correction generalized estimation equation. It can estimate the index parameters and linear parameters in turn, and can obtain the asymptotic normality of the estimator. The Monte Carlo simulation results show that the algorithm has good robustness. References [1] Shakhawat Hossain and Le An Lac. Optimal shrinkage estimations in partially linear single-index models for binary longitudinal data[J]. TEST, 2021, 30(4) : 1-25. [2] Quan Cai and Suojin Wang. Inferences with generalized partially linear single-index models for longitaudinal dta[J]. Journal of Statistical Planning and Inference, 2018, 200 : 146- 160. [3] Gaorong Li and Peng Lai and Heng Lian. Variable selection and estimation for partially linear single-index models with longitudinal data[J]. Statistics and Computing, 2015,25(3) : 579-593. [4] Peng Lai and Gaorong Li and Heng Lian. Quadratic inference functions for partially linear single-index models with longitudinal data[J]. Journal of Multivariate Analysis, 2013, 118 : 115-127. [5] Zeger S L, Diggle P J. Semiparametric models for longitudinal data with application to CD4 cell numbers in HIV seroconverters. Biometrics, 1994, 50: 689–699. [6] Yu Ying Jiang. Empirical Likelihood Inference for a Partially Linear Model under Longitudinal Data[J]. Applied Mechanics and Materials, 2013, 2545(353-356) : 3355-3358. [7] Hongmei Lin et al. A new local estimation method for single index models for longitudinal data[J]. Journal of Nonparametric Statistics, 2016, 28(3) : 644-658. [8] Peng Lai and Gaorong Li and Heng Lian. Semiparametric estimation of fixed effects panel data single-index model[J]. Statistics and Probability Letters, 2013, 83(6) : 1595-1602. [9] Ruiqin Tian and Liugen Xue. Generalized empirical likelihood inference in partial linear regression model for longitudinal data[J]. Statistics, 2017, 51(5) : 988-1005. [10] Xue L G, Zhu L. Empirical likelihood for single-index models. J Multivariate Anal, 2006, 97: 1295–1312. [11] Chen J, Li D, Liang H, et al. Semiparametric GEE analysis in partially linear single-index models for longitudinal data. Ann Statist, 2015, 43: 1682–1715. [12] Pang Z, Xue L. Estimation for the single-index models with random effects. Comput Statist Data Anal, 2012, 56: 1837– 1853.