Microsoft Word - 7-王永红.doc Advances in Systems Science and Applications (2010), Vol.10, No.1 ISSN 1078-6236 International Institute for General Systems Studies, Inc. 41-47 Multidimensional Structural Regression Model for Causal Inference under Strongly Ignorable Treatment Assignment∗ Yonghong Wang School of Mathematics and Computer, Harbin University, Harbin 150086, China Email:wangyonghong418@163.com Abstract In the study of epidemiology aetiology, we usually cannot measure exposed effect relative to an individual, but under some assumptions, we approximately replace the exposed effect by estimator of population average causal effect. A multidimensional structural regression model for causal inference is established to estimate the population average treatment effect under strongly ignorable treatment assignment. Under the normal distribution, the maximum likelihood estimator for population average treatment effect is proved to be consistent, unbiased and asymptotically normal. Keywords Strongly ignorable treatment assignment Causal inference Population average treatment effect Multidimensional structural regression model Maximum likelihood estimator 1. Introduction The objective of epidemiology study is to search aetiology, and to measure its causal effect in the light of quantity, thereby prevent from the occurrence of disease[1]. The case -control study, which attempts to find a contrasted or exposed group which can be compared with the treated or exposed one, is the common method that we employ in the research of epidemiology aetiology. Except the difference of being exposed and unexposed, the ideal treated group and the contrasted group are expected to share the rest of the other features, which serves as the guiding principle in the epidemiology research[2-3]. The dummy truth model laid a theoretical foundation for us to demonstrate this principle[4]. Let U be the study population, and denote a generic individual in U by u U∈ . A variable E is defined on each u U∈ so that ( ) 1E u = if u is exposed to the casual agent of interest and ( ) 0E u = if u is not so exposed. )(1 uY and )(0 uY indicate the diseased status of u as in the exposed and non-exposed cases respectively. Thus we can define the exposed causal effect of the individual u as )(1 uY — )(0 uY .As we know, in any real epdidemilogic study, one does not observe )(1 uY and )(0 uY simultaneously. Therefore the casual effect of the exposure to the individual is not attainable. However, in some circumstances, we can get the population average causal effect, which can be defined as 1 0 E(Y (u)-Y (u)) , while )(⋅E denotes expection or population average (over U ). Under the circumstances of random tests, for the fact that the independence of the exposed E and other variants, the result should be E (Y1,Y0),which indicates that the exposed E is independent of the random Vector(Y1,Y0). Thus the population average causal effect is shown as )( 01 YYE − = )()( 01 YEYE − = )0()1( 01 =−= EYEEYE .In other words, the causal effect can be estimated by the difference of the exposed population average diseased condition the non-exposed population average diseased condition[5]. The epidemiology, however, is an ∗This work has been supported by the Science Development Foundation of Harbin University(HXKQ200701). Wang: Multidimensional Structural Regression Model for Causal Inference under Strongly Ignorable Treatment Assignment 42 observatory science in nature, so the above formula is not to be established which necessitates the assumption of the state of being ignorable[5-6] . Definition 1.1 Given the covariant X , the exposed E is called being strongly ignorable. If (1)(Y1,Y0) XE , i.e. the observed covariant X is given, the response variant is independent of the exposed E. (1) (2) 1)Pr(0 <=< XeE , i.e. for each and every X, they may receive various treatment. When the E is strongly ignorable, by getting the matching-differencing-averaging between the exposed and non-exposed groups, we can get the unbiased estimate of population average causal effect[7]. But the above method is complicated in calculation and the additional sparse-data problems, excess covariant, though few, may arise. In addition, the method of ranking the individuals with the estimated propensity score is an advisable way to solve the problem, but the calculation is still very complicated[8]. [9] Under strongly ignorable treatment assignment, a structural regression model for causal inference is established to estimates the population average treatment effect. They are the series of results achieved. with the response variant being one dimensional. The paper will discuss the inference problem with multi-dimensional response variants. 2. Multidimensional Structural Regression Model Definition 2.1 In Condition of Strongly Ignorable Treatment Assignment, denote multidimensional Structural Regression model as follows: ⎪⎩ ⎪ ⎨ ⎧ Σ ++= tjtjttjxxtj tjtjtttj eXVNeNX eXY oft independen is ),,0(~),,(~ μ βα (2) Here t=0,1 means two treatment; tjY is the random vector of order 1×p ,indicated the response variant of the jth observation for the tth treatment; tjX is the covariant vector of order 1×k ; tje is the random error vector of order 1×p ; tα is the parameter vector of order 1×p ; tβ is the parameter matrix of order kp× ; xμ is the parameter vector of order 1×k ;The covariant matrix tV和xΣ of tjX and tje are both positive definite. Under model (2) ⎟⎟ ⎠ ⎞ ⎜⎜ ⎝ ⎛ + =⎟⎟ ⎠ ⎞ ⎜⎜ ⎝ ⎛ xtt x tj tj Y X E μβα μ , tA ' ' tj x x t tj t x t x t t X Cov Y V β β β β ⎛ ⎞⎛ ⎞ Σ Σ = ⎜ ⎟⎜ ⎟ Σ Σ +⎝ ⎠ ⎝ ⎠ We can easily find txt VA ⋅Σ= ;because xΣ and tV are both positive definite, we can know 0≠Σ x , 0≠tV ;so 0≠tA ;therefore 1− tA is valid, and we denote it as Bt .again based on the principle of block matrix inversion, we can get Bt= ⎟⎟ ⎠ ⎞ ⎜⎜ ⎝ ⎛ −+Σ −− −−− 11 1'1'1 ttt tttttx VV VV β βββ So ⎟⎟ ⎠ ⎞ ⎜⎜ ⎝ ⎛ ⎟⎟ ⎠ ⎞ ⎜⎜ ⎝ ⎛ +⎟⎟ ⎠ ⎞ ⎜⎜ ⎝ ⎛ t xtt x tj tj AN Y X ,~ μβα μ Where: t=0,1 ; j=1,……,nt . We also get observation logarithm likelihood function of the sample ( )tjtj yx , with the Advances in Systems Science and Applications (2010), Vol.10, No.1 43 capacity nt (t=0,1; j=1,……,nt ) ⎟⎟ ⎠ ⎞ ⎜⎜ ⎝ ⎛ −− − ⎟⎟ ⎠ ⎞ ⎜⎜ ⎝ ⎛ −− − +=− ∑∑∑ = == xtttj xtj t t n j xtttj xtj t tt y x B y x AnL t μβα μ μβα μ ' 1 0 1 1 0 lnln2 Furthermore, 1 1 ' 1 ' 1 0 0 1 2 ln (ln ln ) ( 2 tn t x t tj x tj tj x x t t j L n V x x x μ− − = = = − = Σ + + Σ − Σ∑ ∑∑ tjtttjtjttttjxxx yVxxVx 1''1''1' 2 −−− −+Σ+ βββμμ (3) )22 1'1'1'1'' ttttttjtjttjttttj VVyyVyVx ααααβ −−−− +−++ 3. Conclusion 3.1 Two Lemmas In order to get the likelihood estimates of all parameters in model (2) and the population average causal effect, we give the following two lemmas, of which lemma3.1 is shown at [10]. First, suppose X is the matrix of order nm× , )(xf is the real-valued function of matrix X , x xf ∂ ∂ )( indicate real-valued function )(xf settling partial derivation for each element of matrix X . Thus the following lemmas are valid: Lemma 3.1(1) '')( BA X AXBtr = ∂ ∂ (2) '' ' )( XBAAXB X AXBXtr += ∂ ∂ (3) '1 )( ln −= ∂ ∂ X X X Of which A and B are two matrixes that can make matrix calculation with X , ( )tr X means the trace of matrix X . Lemma 3.2 If model (2) is valid, when the covariant matrix and the random error matrix are independent of each other, then ∑ = = tn j tj t t X n X 1 1 and ∑ = −−= t tt n j ttjttj t XX XXXX n M 1 '))((1 ' , ∑ = −−= t tt n j ttjttj t Xe XXee n M 1 '))((1 ' are independent of each other, where ∑ = = tn j tj t t e n e 1 1 . Proof: First to prove tX and ' tt XX M are independent. Wang: Multidimensional Structural Regression Model for Causal Inference under Strongly Ignorable Treatment Assignment 44 Let ⎟ ⎟ ⎟ ⎟ ⎟ ⎠ ⎞ ⎜ ⎜ ⎜ ⎜ ⎜ ⎝ ⎛ = ktntntn kttt kttt t ttt XXX XXX XXX X 21 22221 11211 Set '1 tt X n X = 1 also '' 1(1 ' t t t t XX X n X n M tt −= 11 ' )( t t n X 1 − 11 ' tX ) = '1 t t X n (I t n nt 1 − 11 ' ) tX Which 1 is the column vector of 1×tn . We can do the vector straight operations to matrix tX , then ⊗Σ= xXVecCov ))(( I tn , also, as for tX and each at the element ijm in ' tt XX M follow tX = tn 1 (Ik⊗ 1 ' ) )( tXVec ijm = ')( tXVec [ ⊗+ )( 2 1 ' ijij EE (I t n nt 1 − 11 ' )] )( tXVec Of which⊗ indicates the Kronecker product of matrix, ijE indicates matrix of order kk × with (i,j)th is 1, and the rest is 0. Now we can draw the following conclusion by making use of the characteristic of multidimensional normal random variant and the Kronecker product of the matrix: tn 1 (Ik⊗ 1 ' )( ⊗Σ x I tn )[ ⊗+ )( 2 1 ' ijij EE (I t n nt 1 − 11 ' )] = t ijijx n EE 1[)]( 2 1[ ' ⊗+Σ 1 ' (I t n nt 1 − 11 ' )]=0 Thus we can prove ijm and tX are independent, and it’s valid to any i,j=1,……,k. so we have proved ∑ = = tn j tj t t X n X 1 1 and ∑ = −−= t tt n j ttjttj t XX XXXX n M 1 '))((1 ' are independent of each other. Suppose ⎟ ⎟ ⎟ ⎟ ⎟ ⎠ ⎞ ⎜ ⎜ ⎜ ⎜ ⎜ ⎝ ⎛ = ptntntn pttt pttt t ttt eee eee eee e 21 22221 11211 ),( ttt eXZ = ,in the same way , and by mutual independence of tX and te , we can prove tX and ' tt Xe M are also independent of each other. Lemma3.2 is proved! Advances in Systems Science and Applications (2010), Vol.10, No.1 45 3.2 Estimate of Parameter Theorem 3.1 Suppose model(2)is valid, we can get the maximum likelihood estimates of all the parameter: txxxytt xMMy tttt 1)(ˆ '' −−=α ; 1)(ˆ '' −= tttt xxxyt MMβ ; ∑∑ = =+ = 1 0 110 1ˆ t n j tjx t x nn μ ; ' 1 0 110 )ˆ)(ˆ(1ˆ xtjx t n j tjx xx nn t μμ −− + =Σ ∑∑ = = ; ∑ = −−−−= tn j tjtttjtjtttj t t xyxy n V 1 ')ˆˆ)(ˆˆ(1ˆ βαβα Of which ∑ = = tn j tj t t y n y 1 1 , ∑ = −−= t tt n j ttjttj t xy xxyy n M 1 '))((1 ' ; The form of tx and ' tt xx M is shown in lemma3.2. Proof: First, we get the score function of tα : ∑ = −−− +− ∂ ∂ = ∂ −∂ tn j ttttttjttttj tt VVyVxL 1 1'1'1'' )22()ln2( ααααβ αα =∑ = −−− ∂ ∂ + ∂ ∂ − ∂ ∂tn j ttt t tttj t ttttj t VtrVytrVxtr 1 1'1'1'' )]()(2)(2[ αα α α α αβ α = )22( 11 1 tttjtt n j VxV t αβ −− = −∑ =0 We can get tttt xy βα −= . And the score function of tβ is as follows: 0)222()ln2( 1 '1'1'1 =+−= ∂ −∂ ∑ = −−− tn j tjtttjtjttjtjtt t xVxyVxxVL αβ β Simplified as 0)( 1 ''' =+−∑ = tn j tjttjtjtjtjt xxyxx αβ Fit tβ into above formula of tα , we set txxxytt xMMy tttt 1)(ˆ '' −−=α . In the same way, we can get the maximum likelihood estimates of the other three parameters. Theorem3.1 is proved! In fact, as for the multidimensional structural regression model with the response variants as shown in model (2), we can see from the above demonstration that the maximum likelihood estimates of parameter tα and tβ do not depend on the selection of the covariant matrix and random error matrix. Moreover, we can also see that if we want to estimate the parameter that measure relationship between the covariant and the jth(j=1,……,nt)response variable, we should only note the jth response variable. i.e. the model(2) can be divided into p one dimensional models, which is evident in Jinhua’s paper(2000). 3.3 Population Average Causal Effect Definition3.1 provides the covariant X, if the assumption of Strongly Ignorable condition (1) is fulfilled, then treatment E=1 to E=0 average causal effect ATE is ))()(()()( 0101 xXYEYEEYEYEATE x =−=−= Wang: Multidimensional Structural Regression Model for Causal Inference under Strongly Ignorable Treatment Assignment 46 xx x xxE xXEYExXEYEE μββααβαβα )()()( )),0(),1(( 01010011 01 −+−=−−+= ==−=== Theorem3.2 suppose model (2) is valid, and treatment assignment variant E is strongly ignorable, then (1) The maximum likelihood estimator of population average treatment effect ATE is ATE = )( ˆˆ 01 10 0110 01 XX nn nn YY − + + −− ββ (2) ATE is consistent unbiased estimator of ATE. (3) With the case of the large sample, ETA ˆ is close to normal distribution N(ATE, Γ). Of which ' 0 1 0 1 01 0 1 0 1 ( ) ( )xV V n n n n β β β β− Σ − Γ + + + Proof: (1) with model (2), if treatment assignment variant E is strongly ignorable, then the population average treatment effect ATE= xμββαα )()( 0101 −+− Based on theorem 3.1 and the invariance of maximum likelihood estimator, we can know ATE xμββαα ˆ)ˆˆ()ˆˆ( 0101 −+−= )( ˆˆ 01 10 0110 01 XX nn nn YY − + + −−= ββ (2) Consistency: we know based on the model (2) ttttt eXY ++= βα , t=0,1 Again based on theorem3.1, we get 1)(ˆ '' −= tttt XXXYt MMβ = 1))(( '' −+ tttt XXXet MMβ Based on the law of large numbers and the nature of the random variant order converging by probability, we can know x P XX tt M Σ⎯→⎯' , 0' ⎯→⎯P Xe tt M Then t P t ββ ⎯→⎯ˆ In the same way, x P tX μ⎯→⎯ , xtt P tY μβα +⎯→⎯ So far, again based on the nature of the random variant order converging by probability, we get PATE ATE⎯⎯→ , i.e. ATE is consistent estimate of the population average treatment effect ATE. Unbiased:Through the demonstration of the consistency, we know 1 1 ' )]())((1[ˆ ' − = ∑ −−+= tt t XX n j ttjttj t t MXXee n ββ Whereupon })]())((1[{)ˆ( 1 1 ' ' − = ∑ −−+= tt t XX n j ttjttj t t MXXee n EE ββ = })]())((1[{ 1 1 ' ' ttXX n j ttjttj t xt xXMXXee n EE tt t =−−+ − = ∑β = tttxt xXE ββ ==+ )0( Therefore xtttYE μβα +=)( , so based on lemma3.2, we set Advances in Systems Science and Applications (2010), Vol.10, No.1 47 0 1 1 0 1 0 1 0 0 1 ˆ ˆ ( ) ( ) ( ) ( _ )n nE ATE E Y Y E E X X n n β β+ = − − + = ATEx =−+− μββαα )()( 0101 (3) Asymptotically normality: Based on above demonstration, we know ( )E ATE ATE= Also because of 0 1 1 0 1 0 1 0 0 1 ( ) ( )n nATE Y Y X X n n β β+ ≈ − − − + = ] )( [] )( [ 00 10 010 011 10 011 1 eX nn n eX nn n + + − +−+ + − + ββ α ββ α Moreover x t t n XVar Σ= 1)( , t t t V n eVar 1)( = Therefore 1 1 ' 01012 10 1 11 10 011 1 1)()( )( ) )( ( V nnn n eX nn n Var x +−Σ− + =+ + − + ββββ ββ α 0 0 ' 01012 10 0 00 10 010 0 1)()( )( ) )( ( V nnn n eX nn n Var x +−Σ− + =+ + − + ββββ ββ α So ' 1 0 1 0 1 0 1 0 0 1 1 1 1( ) ( ) ( )xVar ATE V V n n n n β β β β= + + − Σ − Γ + So based on the Central Limited Theorem, we know ~ ( , )ATE AN ATE Γ . Theorem 3.2 is proved! References [1] Rothman,K.J. Modern epidemiology[M]. Boston:Litte, Brown and Company, 1986: 56. [2] Guo,J.H. Causal inference and confounding phenomenon. Journal of Northeast Normal University, 2002, 34: 25-41. [3] Habtemariam,T. and Tameru,B. etc. Epidemiologic modeling of HIV/AIDS: Use of computational models to study the population dynamics of the disease to assess effective intervention strategies for decision-making. Advances in Systems Science and Applications, 2008, 8(1): 35-39. [4] Tao,Q.S. and Li,L.M. Aetiology effect model based on dummy truth theory. Chinese Journal of Epidemiology, 2002, 23(1): 60-62. [5] Rosenbaum,P. and Rubin,D.B. The central role of the propensity score in observational studies for causal effects. Biometrika, 1983, 70: 41-55. [6] Holland,P.W. Confounding in epidemiologic studies. Biometrics, 1989, 45: 1310-1316. [7] Kong,F.L. and Wang,Y.F. Certain Martingale Methods in Parameter Estimation. Advances in Systems Science and Applications, 2004, 4(1): 1-6. [8] Yu,J.W. and Cheng,D.H. etc. Factor analysis methods in the statistical literature of Chinese medicine diabetes research applications. Advances in Systems Science and Applications, 2007, 7(2): 161-166. [9] Jin,H. and Fang,J.Q. Structural regression model for causal inference under strongly ignorable treatment assignment. Journal of south china normal university, 2000, 4: 7-12. [10] Wang,S.G. etc. Theory of linear model[M]. Beijing: Science press, 2004: 44-47.