Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 1s (2025) 1 https://internationalpubls.com Modeling GHQ Total and SDQ Difficulty Score to Extract Incomplete Information on One Using the Other Alka Sabharwal1, Babita Goyal2*, Lalit Mohan Joshi3 1Professor, Department of Statistics, Kirori Mal College, University of Delhi, Delhi-110007 Email: alkasabh@gmail.com 2 Professor, Department of Statistics, Ramjas College, University of Delhi, Delhi-110007 Email: goyalbabita@gmail.com 3Research scholar, Department of Statistics, University of Delhi, Delhi-110007 Email: lalitjstats12@gmail.com *- Corresponding author Article History: Received: 05-08-2024 Revised: 25-09-2024 Accepted: 07-10-2024 Abstract: Background: A natural approach to analyze multidimensional data is use of multivariate statistical analysis. If the number of dimensions/variables is small and variables are correlated, Multivariate Normal Distribution (MVN) is applied frequently. If the distribution of underlying variables is not normal, transformations are applied to convert them to normal variables. Objective: Many a times, information may be missing on one or more dimensions of the data. The study intended to estimate the missing information through the available information. Method: In a series of three independent surveys to examine the psychological health of young adults during COVID-19 period, enrolled in higher educational institutions in India, Strength and Difficulty Questionnaire (SDQ) 17+ extended version was used. In addition, General Health Questionnaire (GHQ) was used in third survey. The data was divided into two datasets; based on third survey and based on first two surveys. MVN was used to estimate GHQ scores through difficulty score dimension of SDQ and vice versa. The model was applied to data of first two surveys to estimate GHQ scores at the time of these surveys. The model was applied on 162 respondents who were common in all the three surveys. Result: The estimated values for the third survey data were consistent with the observed and simulated values. Further it was found that out of 64 respondents with high GHQ scores in third survey, 55 had it during first two surveys also. Conclusion: The results can be extended to estimate any missing information whenever variables are correlated. Keywords: Bivariate normal distributions, General Health Questionnaire, Psychological health, Strength and Difficulty Questionnaire, Weibull distribution. 1 Introduction Any medical data, whether on an individual or on a cohort is multidimensional in nature, where the number of dimensions can be large. To analyze such data, a natural approach is to make use of multivariate statistical analysis. A dimension refers to specific type information collected on a subject, through variables; such as subject dimension includes information about the patient (name, age, gender etc.) Similarly disease dimension included duration of the disease; severity, comorbidities etc. If the number of variables used to collect the information is large, then the commonly used statistical techniques are General Linear Models, Discriminant Analysis, Factor Analysis, both for dimension reduction and analysis of the data (Johnson and Kotz, 1972; Khattree and Naik 2000; Lindsey, 2000; mailto:alkasabh@gmail.com mailto:goyalbabita@gmail.com mailto:lalitjstats12@gmail.com Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 1s (2025) 2 https://internationalpubls.com Chi, 2012). If the number of variables is small (less than four), parametric multivariate methods are frequently applied. One of such methods is use of Multivariate Normal Distribution (MVN). Multivariate normal distribution is used in case when the different variables of the data are correlated. A MVN involves the correlation matrix of the variables along with their marginal distributions. Although the MVN requires the marginal distributions to be normal, this condition is not met in general. Transformations are then applied to convert the non- normal variable(s) to normal variable(s). The most commonly used transformations are the Box-Cox transformation and the power transformations. If the number of variables is two, the MVN distribution reduces to a bivariate normal distribution (BVN). Lipow and Eidemiller applied BVN distribution to study the relationship between stress and strength in the reliability study problem (Lipow and Eidemiller, 1964). Yue applied BVN distributions to solve problems of hydrological engineering design and management (Yue, 1999). Grover et al. applied MVN and BVN distributions to estimate the duration of diabetes on the basis of Low-density Lipoprotein (LDL), Fasting Blood sugar (FBG) and systolic blood pressure (SBP) (Grover et al., 2014). In another study, Grover et al. applied MVN for estimating the length of stay (LOS) in the hospital, duration of disease and severity of the disease on a group of 146 inpatients diagnosed with mental and behavioural problems (Grover et al., 2015). Psychological health of an individual is essentially assessed with the help of questionnaires which are multidimensional in nature. Sometimes more than one questionnaire is needed to have a comprehensive view of the psychological health of the individual. Purpose of using more than one questionnaire is to cross-validate the data collected through the different questionnaires as well as to complement the missing information, if any. In fact, a judicious choice of questionnaires not only ensures the consistency in the responses but also can be used to overcome the shortcoming of any questionnaire. Questionnaires are designed keeping in mind the target group under observation as well as the purpose of study/observation. In this study, the authors used the General Health Questionnaire (GHQ), and the Strength and Difficulties Questionnaire (SDQ) 17+ extended version. The GHQ, developed by Dr. David Goldberg in 1970, is a commonly used, self-administered screening tool designed to detect current state mental disturbances and the health of the last 4-6 weeks (Goldberg and Hillier, 1979). It consists of 12 items that assess various aspects of an individual's mental health, including mood, anxiety, and social functioning. The SDQ, developed by Robert Goodman in 1997, is used to assess strength and difficulties among children and young adults (Goodman, 1997). It has two versions: the basic version and the extended version. A series of three independent surveys was conducted during COVID-19 pandemic, from May–June 2020 to January–February 2022. Each of the survey was conducted after occurrence of some game- changing event. The first survey was held during May–June 2020, when the lockdown was freshly imposed. The second survey was held in October 2020–February 2021, when the first COVID-19 wave had almost subsided. However, soon after the deadly ‘delta’ wave had struck, causing a huge havoc on the society. The large number of morbidities and mortality severely affected the mental health of people at large. In this light, we conducted the third survey, during the months of January–February 2022. While all the three surveys were conducted using the SDQ in order to measure the psychological health of the young adults, in the third survey we used the GHQ-12 additionally, a questionnaire which is used to measure the recent (up to 4 weeks) psychological health. Whereas SDQ is students’ based questionnaire, GHQ-12 is widely used for every age group. This study was initiated on the basis of the data collected during the third survey as this data contained information collected through SDQ as well as GHQ-12. The objective of this study was to examine if it is possible to estimate one score from the other through appropriate modeling. For this objective, we intended to obtain a relationship (if any) between the information obtained through the two Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 1s (2025) 3 https://internationalpubls.com questionnaires (Difficulty score of SDQ and total score of GHQ-12, based on Likert scale 0-1-2-3). In order to achieve this objective, we selected appropriate distributions (on the basis of AIC and BIC criteria) to the two components of the data. The parameters of the selected distributions were estimated through the method of Maximum Likelihood Estimation (MLE). Through power transformations, both the components were transformed to follow normal distributions. Mardia test was applied to test the multivariate normality. On the basis of the outcome of the test, a bivariate normal distribution was fitted using the two correlated normal variables. Using the parameters of the transformed distributions (fitted to the Difficulty score of SDQ and total score of GHQ-12) and the correlation between them, we simulated 10000 values each of (i) Marginal distribution of the Difficulty score (ii) Marginal distribution of GHQ-12 total (iii) Bivariate normal distribution We used the data from survey 3 (Dataset_1) in order to estimate Difficulty score of SDQ through total score of GHQ and vice versa. The results were validated through simulation. The model was then used on the data of the first two surveys (Dataset_2, containing information on Difficulty score only) to estimate the GHQ-12 scores. The model was also validated by applying it on 162 respondents who had participated in all the three surveys. Novelty of the study is establishment of the link between two different questionnaires, used under different circumstances through modeling. The results can be used to complete any missing piece of information. To the best of our knowledge, this is the first study to make use of real life data to link two different questionnaires. The rest of this paper along with the introduction is as follows. In section 2, the materials and developed model are discussed. In section 3, the results are described. The paper concludes with discussion in section 4, discussion and conclusion in section 5. 2 Materials and Methods 2.1 Materials During COVID-19 pandemic times, in a period of almost two years (From May-June 2020 to January- February 2022), three independent surveys were conducted on the young adults studying in higher educational institutions across India, using SDQ 17+ extended version. The number of responses obtained were respectively 1020, 743 and 934 respectively. Although the surveys were held independent of each other, 162 respondents were found to have participated in all the three surveys. In survey 3, GHQ-12 questionnaire was also used along with SDQ 17+ extended version. The scoring methods of SDQ 17+ extended version have been discussed in the earlier literature (Goodman, 1999; Goyal et al., 2023; Sabharwal et al., 2023). The GHQ-12 questionnaire contains 12 items which are evaluated using the Likert (0-1-2-3/ 0-0-1-1) scale. The scores are then added to obtain the total GHQ scores. The scores lie in the range 0-36/0-12. Higher scores indicate a higher level of psychological distress or impaired mental well-being, although there are no clear-cut categories describing the severity of problems as is in case of SDQ scores (Goldberg, 1979; Goldberg et al., 1997; Anjara S. et al., 2020). The inclusion criteria for the study were the respondents enrolled in higher educational institutions and were participating willingly. Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 1s (2025) 4 https://internationalpubls.com 2.2 Methods 2.2.1 Bivariate Normal Distribution (BVN) If two random variables X and Y are following normal distributions 2~ ( , )x xX N   and 2~ ( , )y yY N   respectively, with correlation coefficient  , the joint probability density function (pdf) of x and y is given by (Johnson and Kotz, 1972), 22 2 2 1 exp 2 2(1 ) ( , ) 2 1 y yx x x y x y x y y yx x f x y                  − −   − −−   + −           −          = − (1) ( , )x y−    , 1 1−   , ( , ) 0x y   The conditional expectation for BVN distribution is E( | ) ( ) xy x y x x y y     = − − (2) where, xy is covariance between x ; xy−    . 2.2.2 Distribution of Variables The following distributions namely, Normal, Gamma, Weibull, Log-normal, and Exponential, were fitted to the data and the best fitted distribution was selected on the basis of minimum AIC and BIC values. Table 1 presents the pdf of all the fitted distributions in this study. Table 1 The fitted distributions for present study Name of distributions Probability density function (pdf) Name of parameters Ranges Normal 2 1 1 ( , , ) exp 22 x f x      −  = −       = Mean and  = Standard deviation ( , ) , 0x  −     Gamma ( / ) ( 1) ( , , ) xe x f x        − − =   = Shape parameter and = Scale parameter ( , , ) 0x    Weibull 1 ( , , ) exp( ( / ) ) k kk x f x k x    −   = −    k = Scale parameter and = Shape parameter 0x  Log-normal 2 2 1 (ln ) ( , , ) exp 22 x f x x      − = −     = Mean log and 0x  Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 1s (2025) 5 https://internationalpubls.com  = Standard deviation log Exponential ( , ) xf x e =  = Rate parameter. 0x 2.2.3 Transformations Transformations are a group of statistical techniques, used to convert the data to a form which can be dealt with available standard methods. The most commonly used transformations are the square root, the logarithms, and the reciprocal (http://www- users.york.ac.uk/~mb55/msc/clinbio/week5/transfm_gif.pdf). For the bivariate data used in this study, the assumed distribution is a bivariate normal distribution (BVN) for which the marginal distributions should also be normal. Power transformations have been applied to the selected distributions to convert them to a normal form. The form of power transformation is (http://www.statsref.com/HTML/index.html?freeman-tukey.html). ( ) , 0, 0, 0z x a a x = +    (3) where a is an optional constant. 2.2.4 Mardia Test (Mardia K., 1970; Von Eye and Bogat, 2004) Mardia test is used to examine the multivariate normality of a data. It makes use of skewness and kurtosis measurements, given by 3 1, 2 1 1 1 ˆ n n p ij i j t n  = = =  and 2 2, 2 1 1 ˆ n p ii i t n  = =  where, X1, X2, …, Xn are a vector of size 1 × p, ' 1( ) ( )ij i jt x x S x x−= − − ' 1 1 ( )( ) n i ii S x x x x n =  = − −   and 1 1 n i i x x n = =  and, p is the number of variables. The test statistics for skewness and kurtosis are, 1, 2 1 ( 1)( 2)/6 ˆ ~ 6 p p p p approx n   + += and ( ) ( ) ( ) 2, 2 ˆ ( 2) ~ 0,1 8 ( 2) p asymp p p N p p n   − + = + respectively. 2.2.5 Criteria for Model Selection The most popular information criteria for choosing a model are the Akaike Information Criterion (AIC) and the Bayesian Information Criterion (BIC), given by: 2 log( ) 2AIC L p= − + (4) 2log( ) log( ).BIC L n p= − + (5) Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 1s (2025) 6 https://internationalpubls.com where, L is the likelihood under the fitted model, p is the number of parameters, and n is number of total observations. The smallest value of AIC or BIC give best fit model for data (Kuha, 2004; Bradman M. et al., 2003). 2.2.6 Algorithm used in the Study The following flow chart describes the algorithm used in this study. Fig. 1 Algorithm of the modeling and its application 3 Results 3.1 Data description Table 2 presents the descriptive statistics of 934 respondents for Dataset_1(Survey 3 data). The mean and standard deviation value of GHQ total is 16.82 and 7.98 respectively. The mean and standard deviation value Difficulty score is 15.11 and 5.90 respectively. The maximum value of Difficulty score is 31 and that of GHQ total is 36. Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 1s (2025) 7 https://internationalpubls.com Table 2 Descriptive statistics of variables: GHQ total and Difficulty score for Dataset_1(Survey 3 data) Statistic GHQ total Difficulty score Total 934 934 Minimum 0 1 Maximum 36 31 Range 36 30 Mean 16.82 15.11 Standard deviation 7.98 5.90 Table 3 presents the descriptive statistics of 1020 and 743 respondents giving minimum; maximum, range, mean, and standard deviation for Dataset_2 (survey 1 and survey 2 data). The mean value of Difficulty score is 13.64 and 12.73 for 1st and 2nd survey respectively. The maximum value of Difficulty score for 1st and 2nd survey is 31 and 29 respectively. Table 3 Descriptive statistics variable; Difficulty score for Dataset_2 (survey 1 and survey 2 data) Statistic Difficulty score for 1st survey Difficulty score for 2nd survey Total 1020 743 Minimum 2 2 Maximum 31 29 Range 29 27 Mean 13.64 12.73 Standard deviation 5.35 5.17 3.2 Distribution Selection for GHQ total and Difficulty scores The Table 4 below presents the AIC and BIC values of the fitted distributions on Dataset_1. On the basis of the least AIC and BIC criteria, Weibull distribution (with computed parameters) has been found to be the most appropriate fir for both the variables. For GHQ total, Weibull distribution is obtained with parameters, the shape parameter ( k̂ ) is 2.266015 and the scale parameter ( ̂ ) is 19.06131. For the Difficulty score, the Weibull distribution has the parameter shape ( k̂ ) 2.790381 and the scale ( ̂ ) 16.98565. Table 4 AIC and BIC values of different distributions for GHQ total and Difficulty score Variable Distribution AIC values BIC values Selected distribution MLE of the parameters GHQ Total Normal 6514.19 6523.87 Weibull k̂ =2.266015 Gamma 6488.11 6497.79 Weibull 6447.84 6457.51 Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 1s (2025) 8 https://internationalpubls.com Lognormal 6592.33 6602.005 ̂ = 19.06131 Exponential 7131.75 7136.59 Difficulty score Normal 5970.19 5979.87 Weibull k̂ =2.790381 ̂ = 16.98565 Gamma 6006.99 6016.67 Weibull 5945.73 5955.41 Lognormal 6117.06 6126.74 Exponential 6942.85 6947.69 3.3 Transformations Applied on GHQ total and Difficulty score The square root transformation has been applied on GHQ total to obtain an approximately normal *GHQ total and the fourth root transformation has been applied on Difficulty score to obtain an approximately normal *Diff score. The normality of the transformed variables *GHQ total and *Diff score has been justified by plotting the pdfs and the Q-Q plots of the transformed variables in Fig. 2(a), 2(b), 3(a) and 3(b). Fig. 2(a) Normal approximation of *GHQ total Fig. 2(b) QQ-plot for *GHQ total Fig. 3(a) Normal approximation of *Diff score Fig. 3(b) QQ-plot for *Diff score Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 1s (2025) 9 https://internationalpubls.com 3.4 Checking Multivariate Normality of *GHQ total and *Diff score Mardia test has been applied on *GHQ total and *Diff score to test the multivariate normality. Skewness and kurtosis of the bivariate data are 1 =2.4134 (at 4 degrees of freedom with p-value = 0.6601 > 0.05), and 2 = -0.3640 (with p-value =0.7158 > 0.05) respectively. Hence, the *GHQ total and *Diff score jointly follow BVN (3.971, 1.939, 1.05598, 0.04516, 0.49997). 3.5 Estimating *GHQ total through *Diff score and vice versa using Dataset_1 through BVN The joint distribution of variables *GHQ total and *Diff score is BVN, given by * total * total * total,* * * total,* * * total ~ , * GHQ GHQ GHQ Diff score Diff score GHQ Diff score Diff score GHQ BVN Diff score              =  =                (6) 3.5.1 Generating the BVN population corresponding to transformed variables *GHQ total and *Diff score through simulation Equation (6) is valid for population parameters. In order to generate the BVN population parameters, we conducted a simulation study as follows: 1. Distributions of GHQ total and Difficulty score were selected on the basis of minimum AIC and BIC value. 2. 10000 values of GHQ total were generated using the selected Weibull (2.2660,19.0613) distribution. 3. 10000 values of Difficulty score were generated using the selected Weibull (2.7903, 16.9856) distribution. 4. We applied the same power transformation as were applied on the real data to obtain the simulated transformed normal variables (sim GHQ total and sim Diff score). 5. The parameters of the transformed variables along with their correlation coefficient were used to generate the bivariate normal population of size 10,000. 6. The values were put in equation (6). 3.5.2 Comparison of data based Mean Difficulty score with estimated Difficulty score as obtained through the BVN model The following procedure is adopted to obtain the mean Difficulty score for various categories of GHQ total: 1. The intervals of *GHQ total (after transformation) range is defined as 1 < GHQ ≥ 2, 2 < GHQ ≥ 3, 3 < GHQ ≥ 4, 4 < GHQ ≥ 5, and 5 < GHQ ≥ 6; identified as GHQi; i = 1,2,3,4,5 respectively. 2. Take the first range of GHQ. Corresponding to this range, compute the mean GHQ from Dataset_1. Fixing this range of GHQ, compute the mean Difficulty score i.e., by taking only those Difficulty scores, corresponding to which GHQ lies in this range. 3. Repeat (2.) for all the ranges of GHQ, thus yielding all the mean GHQ totals and the corresponding mean Difficulty scores. 4. The same procedure (for the same GHQ total ranges) is applied for the simulated data. Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 32 No. 1s (2025) 10 https://internationalpubls.com 5. Next, taking GHQ total from the sample data, the conditional expectation of Difficulty score for each GHQ category (GHQi, i = 1,2,3,4,5) has been computed using equation (7) where all the other parameters are population based (simulated data) ( ) ,, . . , . , . 1 ( . | ) ii i Diff i GHQ Diff i GHQ i Diff Diff GHQ GHQ     = − − (7) 6. The observed mean values of Diff score corresponding to each range of GHQ total, is compared with the estimated values obtained through equation (7). 7. Step (6) is used again after retransforming the transformed variables to the original variables. The results are presented in Table 5 below: Table 5 Estimated mean *Diff score given *GHQ total for 934 respondents for different interval of *GHQ total using a generated random sample of size 10000 for BVN distribution Interval Real Data Simulated Data E[*Diff. | *GHQ] Observed difficulty score Estimated difficulty score Mean *GHQ Mean *Diff. Mean *GHQ Mean *Diff 1