EUROPEAN JOURNAL OF PURE AND APPLIED MATHEMATICS Vol. 10, No. 1, 2017, 133-156 ISSN 1307-5543 – www.ejpam.com Published by New York Business Global Sir Clive W.J. Granger Memorial Special Issue on Econometrics Sir Clive W.J. Granger Model Selection Jennifer L. Castle Magdalen College and Institute for New Economic Thinking at the Oxford Martin School, University of Oxford, UK. Abstract. Clive Granger proposed thick modelling as an alternative to selecting a unique model based on a given criterion, or thin modelling. This stemmed from his research on forecast com- bination and portfolio selection in which using just the best asset or forecast can be suboptimal in many settings. This paper proposes to integrate thick modelling into the general-to-specific model selection literature, yielding the benefits of selecting a set of well-specified encompassing models while taking seriously Granger’s critique of model selection. The paper argues that model uncertainty is addressed by applying selection to narrow down the class of models followed by pooling across the retained set of close specifications. An example using artificial data illustrates the approach. 2010 Mathematics Subject Classifications: 62F07, 62P20, 91B84 Key Words and Phrases: Clive Granger, model selection, thick modelling, model averaging. 1. Introduction Economic modelling requires a synthesis of theory and empirics, typically with varying weights placed on each depending on the researcher’s beliefs or preferences. Granger was clear that one of the main objectives of empirical modelling was to inform the decision making process. He was relatively agnostic about the preferred approach to model building as long as the models were carefully evaluated and validated using actual data, and the models were useful in addressing the research question asked, specifically in terms of the quality of decisions that are made based on the models. Model selection is inevitable in this framework. Imposing theory models with no testing would fail Granger’s criteria, and any other form of model building must necessitate model selection. In Granger’s Marshall Lectures [24] he explains how he sees model selection: Email address: jennifer.castle@magd.ox.ac.uk http://www.ejpam.com 133 c© 2017 EJPAM All rights reserved. J. Castle / Eur. J. Pure Appl. Math, 10 (1) (2017), 133-156 134 A sculptor once said that the way he viewed his art was that he took a large block of stone and just chipped it away until he revealed the sculpture that was hidden inside it. Some empirical modelers view their task in a similar fashion starting with a mass of data and slowly discarding them to get at a correct representation. My perception is quite different. I think of a modeler as start- ing with some disparate pieces some wood, a few bricks, some nails, and so forth - and attempting to build an object for which he (or she) has only a very inadequate plan, or theory. The modeler can look at related constructs and can use institutional information and will eventually arrive at an approximation of the object that they are trying to represent, perhaps after several attempts. Model building will be a team effort with inputs from theorists, econometri- cians, local statisticians familiar with the data, and economists aware of local facts or relevant institutional constraints. The above quote can be misinterpreted by those who are not familiar with the general-to- specific (Gets) model selection literature and perceive the approach to consist of ‘starting with a mass of data and slowly discarding them to get at a correct representation’. In fact, the approach is much more in line with Granger’s view of model building, whereby theory, past evidence and institutional knowledge all inform the initial specification of the general unrestricted model (GUM) but rigorous evaluation and validation is conducted to reduce the model by eliminating those effects that are statistically irrelevant. The important distinction is in the specification of the GUM – without careful thought informed by the objective of the modelling exercise, theory and institutional knowledge, and an understanding of the data and its limitations, one is in danger of data mining. The edited volume by Granger (1990) (see [23]) provides a collection of critiques. Granger’s interest in model selection came later in his research career, see Hendry (2010) [37]. He showed remarkable foresight in Granger, King and White (1995) (see [31]) where he argues that model selection procedures were preferable to hypothesis testing when specifying empirical models, thereby anticipating developments in Gets modelling such as Hoover and Perez (1999) [51] and Hendry and Krolzig (1999) [42]. The general-to-specific literature, often termed the LSE approach due to its proponents at the London School of Economics in the 1960s and 1970s had been developing for some time, see, e.g. Pagan (1987), Phillips (1988), Mizon (1995) and Hendry (2003) ([60],[63],[58],[35]), but were only automated later. Granger, in an interview with David Hendry explored the Gets approach in detail in Granger and Hendry (2005) (see [29]). Although Granger’s contributions to the field of model selection were not as prolific as his contributions to other fields of econometrics demonstrated by the papers in this special issue, he clearly had an interest in model selection as seen in Granger and Hendry (2005) [29]. One of the contributions of this paper is to consider how model selection performs under misspecification, which was a question that Granger highlighted in [29]. Granger coined the phrase ‘thick modelling’ to argue that model selection need not focus on selecting just one model specification from the infinite range of possibilities, which would be defined as the ‘best’ model based on some pre-specified criterion. He proposed keeping all alternative close specifications that are validated on empirical criteria J. Castle / Eur. J. Pure Appl. Math, 10 (1) (2017), 133-156 135 and pool the resulting outcomes, be they parameter estimates, impulse responses, policy simulations, hypothesis tests or forecasts, see Granger and Jeon (2004) and Granger (2005) [30, 26]. Given Granger’s interest in portfolio theory (Granger, 2005, [27]) this is a natural thing to do, and corresponds to his pragmatic approach to empirical modelling discussed in Granger (2009) (see [28]). It also ties with his research on forecasting and forecast combinations stemming from Bates and Granger (1969) (see [5]). The close correspondence between the views on model selection of Granger’s thick modelling and Hendry’s Gets modelling, and their many personal and professional inter- actions on the subject, suggest that the two approaches can be integrated. This paper investigates the approach to model selection of undertaking general-to-specific selection and then applying thick modelling techniques of averaging across the selected models, where one or more models can be retained. We find that thick modelling does not resolve problems of misspecification, but can capture model uncertainty reflected by retaining different model specifications commencing from the same starting point due to searching many reduction paths. Thick modelling across well-specified models captures the relevant model uncertainty implied by model selection. Section 2 outlines the selection approach implemented by Autometrics and discusses how thick modelling can be implemented, sec- tion 3 notes the literature in which thick modelling is employed, section 4 demonstrates the approach with an artificial data series and section 5 concludes. 2. Model selection, model averaging and thick modelling Commencing from an information set, model selection can be interpreted as pooling the available information set and selecting possible model specifications therefrom. Thick modelling will retain a set of specifications and thin modelling will retain a single specifica- tion. Model averaging pools the model specifications resulting from subsets of information. The distinction between pooling information and pooling models has implications for con- gruency, which will now be discussed. 2.1. Model selection Granger emphasized the importance of establishing the purpose of the econometric modelling exercise. Let us assume that the objective is to identify a model(s) that matches the unknown Data Generation Process (DGP) in all measured aspects. This introduces the notion of congruence which relates to the model being statistically well-specified, see e.g., Bontemps and Mizon (2003) [6]. Congruence requires that the empirical model is statistically ‘close’ to the evidence. The Theory of Reduction establishes the properties of the unknown local data generating process (LDGP) which delivers a mean-zero, ho- moskedastic innovation process that the model aims to capture. Hendry (1995, ch.9) (see [34]) provides extensive discussion.∗ A further step in the Gets literature is that the model ∗The Theory of Reduction commences from an unmanageably large DGP denoted by Du ( U1 T |U0, ψ 1 T ) , with ψ1 T ∈ Ψ ⊆ RkT , where U1 T = (u1, . . . ,uT ) is the full sample vector of random variables defined on probability space (Ω,F ,P). A series of reduction steps are applied to the DGP, including sequential J. Castle / Eur. J. Pure Appl. Math, 10 (1) (2017), 133-156 136 should encompass every other model that is a valid restriction of the general unrestricted model. Encompassing refers to the ability of a model to explain the results of rival models and hence make them redundant, enabling selection between two or more congruent mod- els. Mutually encompassing models are able to explain the results of each other’s model, so the models cannot be ranked. The LDGP will be congruent and will encompass all other models, so models that are non-congruent or non-encompassing fail to model the LDGP. Encompassing has been extensively discussed in, inter alia Mizon and Richard (1986) [59], Hendry and Richard (1989) [45], Mizon (1994) [57] and Hendry (1995, ch.14) [34]. The most recent generation of automated Gets software, Autometrics (see Doornik, 2009, [19] and Hendry and Doornik, 2014, [39]), uses a tree search to eliminate irrelevant variables, with various methods for speeding up the full multipath search. Diagnostic checks are used to ensure the resulting terminal models are congruent. Encompassing tests are implemented to ensure the terminal models encompass the GUM so there is no significant loss of information by undertaking the reduction. They are also performed against the union of the terminal models. This may result in a single unique terminal model being retained, but in general there may be multiple terminal models, all of which are valid reductions of the GUM and mutually encompass each other. In practice, the automated algorithm selects one model based on the Schwarz (1978) (see [64]) information criterion, but a range of other methods to select between congruent encompassing models could be introduced at this stage, or the user may have a preference for a particular terminal model on theory or aesthetic grounds. One of the criticisms of such an approach is that there is no accounting for ‘model’ or ‘specification’ uncertainty. It is argued that the standard errors of the selected terminal model capture the estimation uncertainty due to sampling, but not the uncertainty of the choice of model selected. In response to this criticism, searching many paths increases the probability of locating the LDGP. However, thick modelling provides a route to capturing model uncertainty due to selection, while retaining the benefits of selection. Model selec- tion is designed to handle more complex problems than just locating relevant variables, by jointly including variables, dynamics, outliers and structural breaks, non-stationarities and non-linearities. Castle, Doornik and Hendry (2011) (see [9]) provide simulation evi- dence establishing that Autometrics is able to recover the LDGP almost as often as when commencing from the LDGP itself, and hence the costs of searching over a more general initial specification are small relative to the costs of inference directly conducted on the factorization to obtain an innovation process, marginalization and conditional factorization (to discard the marginal distribution if weak exogeneity is satisfied), and further transformations including lag truncation, functional form, constancy, etc, to result in an LDGP given by A (L) g (yt) = B (L) h (zt) + εt where εt ∼ app Nn (0,Σε) where εt is a mean-zero homoskedastic innovation process with variance Σε and A (L) and B (L) are lag polynomials. . The properties of the LDGP are determined by the reductions applied, so by ensuring there is no loss of information in the reduction process we face the null hypothesis that the LDGP has homoskedastic near normal innovation errors, with weakly exogenous conditioning variables and constant parameters and encompasses the DGP, see Hendry (2009) [36] and Hendry and Doornik (2014, ch.6) [39]. So despite not knowing the DGP or LDGP specification, we can establish the properties that the LDGP should possess. J. Castle / Eur. J. Pure Appl. Math, 10 (1) (2017), 133-156 137 LDGP. Model selection does incur costs—relevant variables can be omitted and irrelevant variables can be incorrectly retained—but these costs are quantifiable using measures of potency and gauge. Potency, which calculates how likely relevant variables are retained (the average non-null rejection frequency for the null hypothesis that a coefficient is equal to zero), should be close to the power of the corresponding test in the LDGP. The nominal null rejection frequency chosen for selection, i.e. the significance level for hypothesis tests of the kind H0 : βj = 0, should match the retention of irrelevant variables (the empirical null rejection frequency), which is called the gauge. These measures refute the criticism by Leamer (1983) (see [55]) that the model selection process is often hidden when the final model is reported. If there are many irrelevant variables, then uncertainty appears to be large as different irrelevant variables are likely to be selected on different draws by chance sampling, commencing from the same information set, but those irrelevant variables will have little impact given that their retention rate is controlled by the nominal significance level. The important aspect of selection is to retain the same set of relevant variables in different draws commencing from the same information set, which will be a function of the non-centralities (population t−statistics) of the relevant variables and the nominal significance level. Forcing the ‘theory-relevant’ variables and implementing selection over only the additional candidate variables will control this cost, see Hendry and Johansen (2015) [40]. 2.2. Model Averaging Model averaging is an alternative to model selection that claims to address model uncertainty. By averaging over 2k possible models within the given model class, where k is the number of potential variables, it is thought that the uncertainty over which model is correct is captured. See Hoeting, Madigan, Raftery and Volinsky (1999) [50] for a Bayesian interpretation and Hansen (2008) [33] for a classical motivation within forecasting. There is a huge literature discussing the appropriate weights for model averaging depending on the purpose of the modelling exercise, for example Akaike (1973), Schwarz (1978) or Hannan and Quinn (1979) (see [2][64][32]) information criteria for in-sample averaging or out-of-sample mean square forecast error criterion if the purpose of the exercise is to forecast. Model averaging is most commonly used in forecasting, where combining forecasts can outperform the individual forecasts if there are offsetting biases or breaks, see Hendry and Clements (2004) [38]. Model or forecast averaging can also help if diversification provides a reduction in variance, see Granger (2002) [25]. However, if proponents of model averaging do wish to account for model uncertainty, then all possible models must be considered rather than a reduced set based on some model selection criterion, unless this selection uncertainty is also accounted for in measuring model uncertainty, which is not typically done in the literature. Many procedures and algorithms have been proposed to narrow down the search space, including Occams razor, branch and bound techniques (see Hocking and Leslie, 1967, [47], LaMotte and Hocking, 1970, [54], and Gatu and Kontoghiorghes, J. Castle / Eur. J. Pure Appl. Math, 10 (1) (2017), 133-156 138 2006, [21]), ridge regression (see Hoerl and Kennard, 1970, [49, 48]), the non-negative garrote (see Breiman, 1995, [7]) and the least absolute shrinkage and selection operator (LASSO: see Tibshirani, 1996, [65]). The cost of averaging over all possible models is that a large subset will necessarily be mis-specified and could contaminate the averaged model. Moreover, the average calculated from different data draws could still differ greatly when regressors are highly collinear and the information set over which models are averaged is not a good representation of the DGP. 2.3. Thick Modelling Granger’s thick modelling can be viewed as an attempt to reconcile the distinct methodologies of model selection and model averaging. The argument that justifies model averaging, namely accounting for model uncertainty, is fallacious—averaging over all pos- sible models of which many must necessarily be non-congruent and therefore lead to biased coefficient estimates cannot be a sensible way to proceed, even if weights are rebased and intercepts are forced, see Hendry and Reade (2008) [44]. Furthermore, model averaging does not help if there are structural breaks in the data, whereas Gets selection seeks well- specified encompassing models of the LDGP and in doing so handles possible structural breaks, which is seen as more important than determining the uncertainty in a badly specified representation. However, discarding potentially useful information in ignored specifications because only one model is selected can also lead to poor outcomes, be they parameter estimates or forecasts. Thick modelling is an approach that keeps all close alternative specifications and pools across these. In the Gets terminology, the retained terminal models are all congruent encompassing models of the LDGP, so pooling across these models will give a thick representation of the LDGP. We can summarise the approaches as the choice of weights on the model combination. If there are k potential variables, there are 2k combinations of subsets of variables, including the null and full set, resulting in 2k possible models. Let M denote the full set of models and Ms the selected set of models, where Ms ⊆ M. Given a criterion cm, the weights are given by: wm = { cm∑ j∈Ms cj m ∈Ms 0 m /∈Ms (1) allowing us to characterize: 1. Model selection (thin modelling in Granger’s terminology): Ms has a single member so wm = 1 for the selected model and 0 otherwise. 2. Unweighted model averaging: Ms =M and cj = 1, ∀j, resulting in wm = 2−k. 3. Thick modelling: assuming L models withinMs, equal weighted averaging overMs with wm = L−1. The criterion for Autometrics is the nominal significance level α which is set by the user. This implies the set of selected models is a function of α,Ms (α), and varying α will J. Castle / Eur. J. Pure Appl. Math, 10 (1) (2017), 133-156 139 324 C.W. Granger, Y. Jeon / Economic Modelling 21 (2004) 323–343 Fig. 1. Thin modeling. Ysg(X,u)q´ where X, Y are, respectively, independent and dependent variables,g( , ) is some. . function, u an unknown parameter(possibly a vector) and ´ is a residual. Given this specification, the only task remaining for the econometrician is the estimation of u and discussion of the properties of this estimate denoted , wheren is theûn sample size. The plot of againstX gives a ‘thin’ curve, as in Fig.ˆw xE Y±X sg(X,u )n 1. This curve has length but no width. The analysis then proceeds by using this model for whatever purpose it was originally constructed: for a policy scenario, for testing a hypothesis, or for forming forecasts, for example. This pristine sequence of events in selecting, estimating, and using a single specification sounds well in a textbook but is not realistic. In practice there is rarely a single specification that is considered; there may be alternative functionsg( , ). . available, alternative forms for the explanatory variable termX, and different ways to estimateu, for example. If each specification is estimated and plotted as before, one will get a ‘thick’ representation as in Fig. 2, where it is expected that the alternatives do not differ greatly from one to another. The present standard practice2 is to choose the ‘best’ specification from those that are available, using some sensible criterion, such as one of the model selection criteria, AIC, BIC, or the maximum likelihood estimator, for example, and so reduce the thick model to a best thin one. This is then used for forecasting or whatever is the desired purpose. However, here any useful information that is in the specifications that are not selected is being ignored. There are good reasons for thinking that this may not be The use of the terms ‘thick’ and ‘thin’ modeling here has little in common with the earlier use of2 these terms by McCloskey(1988)when discussing methodologies. She was writing about methodology being ‘thin’ if it asked ‘thin’ questions; that is, ones that are two narrow or constrained to be of real interest. 325C.W. Granger, Y. Jeon / Economic Modelling 21 (2004) 323–343 Fig. 2. Thick modeling. a good strategy. For example, the combination of forecasts is often a better procedure than using the individually best forecast. Similarly, a portfolio of assets is usually better than investing in a single asset, even if it is better than any other single alternative in comparisons of pairs. Thick modeling consists of using many alternative specifications of similar quality, using each to produce the output required for the purpose of the modeling exercise, such as a set of forecasts, policy scenarios, elasticity estimates, or tests of some hypothesis, and then combine or synthesize the results. One gets a simple form of ‘meta-analysis,’ which is a form of synthesis of research used more in other fields than in economics.3 As a conceptional illustration, consider the estimation of an impulse response of an increase in, say, money supply(M) on unemployment rate(U) in 6 months time. If one builds a vector autoregressive(VAR) model involving m variables, including M andU, with p lags, denote the vector of variables by then the VARXt ¯ is given by X sA (B)X qet p t t ¯ ¯ ¯ ¯ Even though our work is adjacent to meta-analysis, for example as explained in Stanley(2001),3 we do not use it directly. Instead, we are merely discussing reduction(many-to-few) models rather than trying to discuss the relationships to the many, which meta-analysis would do. At the thick stage we only suggest actually combining forecasts, just considering a set of alternative estimators of a parameter in some way. And Leamer(1983) and Leamer and Leonard(1983) adopt the extreme bounds analysis as an approach to the model specification problem. The parameters of interest are estimated and, depending on whether the range of these estimates including a maximum and a minimum effect is narrow enough, the data can be considered to yield useful results. As the referee points out, our approach is similar, but we utilize averaging models instead. Figure 1: Figures from Granger and Jeon (2004) describing thin and thick modelling vary the number of selected models. When α = 1 the GUM will be retained, so Ms will have a single member although no selection has been applied. A tight α will likely lead to a smaller subset of congruent mutually encompassing models retained, so fewer models to average over for the thick modelling. Figure 1, taken from Granger and Jeon (2004) (see [30]), captures the principle of thick modelling. Although Granger did not frame his proposal of thick modelling within Gets selection, the benefits of doing so are twofold. First, Gets aims to select congruent encompassing specifications. Models that are discarded fail this criteria, either because there are irrelevant variables in the model, the model fails to encompass the GUM so information is lost, or the model fails diagnostic tests so cannot be a close approxima- tion to the LDGP whose properties are known even though its specification is unknown. Therefore, averaging over the retained set ensures a well-specified thick model. Second, the intention of thick modelling is to consider alternative specifications of similar quality, which the terminal models after Gets selection must satisfy to be mutually encompassing. The final stage of choosing one terminal model using some pre-specified criteria can be arbitrary and thick modelling will avoid this reliance on the chosen criterion. The advantage of applying the procedure within the Gets framework is that most un- certainty between model specifications will be captured by the range of terminal models retained. A variable with a high non-centrality will most likely be retained in all specifi- cations so the degree of model uncertainty is small, but distinguishing between, say, lag 1 and lag 2 of a persistent variable may be much more difficult, resulting in terminal models that retain one or other of the lag length specifications. The information content of both specifications is close, and therefore averaging across both lags of the variable would be a sensible way to capture the timing uncertainty. A similar consideration applies to vari- ables that have small non-centralities in the LDGP, where a given draw may result in the relevant variable being statistically insignificant, and vice versa, irrelevant variables may chance to have significant selection statistics in the particular data draw. This more closely reflects model uncertainty in the robust model specification sense, where alternative draws for the innovation could result in slightly different model specifications selected. Factors affecting the selected set Ms for thick modelling will include the degree of correlation between regressors and possible misspecification of the GUM, as well as α. J. Castle / Eur. J. Pure Appl. Math, 10 (1) (2017), 133-156 140 High correlations between candidate variables contaminate the ranking of variables based on t-statistics. Assume a GUM which is orthogonal in the sample: yt = N∑ k=1 βkxk,t + εt (2) where T−1 ∑T t=1 xk,txj,t = λkδk,j ∀k, j, where δk,j = 1 if k = j and zero otherwise, with εt ∼ IN[0, σ2ε ], independently of the {xk,t}, and T >> N . The DGP is nested within the GUM, where n ≤ N of the regressors have non-zero βk. In this setting, orthogonality enables 1-cut selection, in which the N sample t2-statistics testing H0: βk = 0 are ordered from largest to smallest and variables are retained if, and only if, their associated t2- statistic is greater than the critical value for a chosen significance level cα: t2(ñ) ≥ c 2 α > t2(ñ+1) (3) where ñ is the cut-off between the number of retained and excluded variables. In this setting, Ms (α) will have a single member and there is no need for thick modelling, other than exploring the consequences of choosing different values of α. If the regressors are correlated, a multipath search must be conducted as the orderings based on t2-statistics no longer represent the significance of the variables in the LDGP. Making a decision to retain/exclude a variable will affect the t-statistics of other variables, thereby changing the ordering. It is the multipath search that can lead to multiple terminal models which are mutually encompassing and this is more likely the higher the degree of correlation. One solution would be to find an orthogonal representation, so transforming an Autore- gressive Distributed Lag (ADL) model to its Equilibrium Correction Mechanism (EqCM) representation, for example, would reduce the correlation structure. But some correlations are unavoidable, for example if a researcher was uncertain if the consumer price index or retail price index was the relevant measure of inflation and therefore included both in the GUM. These are highly correlated measures of inflation and if each was retained in different model specifications, thick modelling would capture this uncertainty. Misspecification of the GUM can also to lead to a larger Ms (α) as various irrelevant variables are retained in an attempt to proxy the variables omitted from the LDGP that are the cause of misspecification. Autometrics (Doornik, 2009, [19]) will still undertake the search procedure even if the GUM fails diagnostic tests, with a process of tightening the significance level of the diagnostic tests if they fail. The diagnostic tests can be user- specified, but for time-series modelling we include tests for autocorrelation, autoregressive conditional heteroskedasticity, normality, functional form and parameter constancy. Auto- metrics will try to restore diagnostic validity along the path searches if the GUM initially fails diagnostic tests. If there are multiple terminal models and some of those pass the diagnostic tests at the original significance level prior to it being reduced, then only those terminals are retained, so thick modelling in this setting would average over the congruent set of terminal models. Castle and Hendry (2011, 2014) (see [12, 14]) discuss the implications of applying model selection to an underspecified GUM, i.e. one that omits relevant variables so the unknown J. Castle / Eur. J. Pure Appl. Math, 10 (1) (2017), 133-156 141 LDGP is not nested. In this setting, there can be no hope of finding a set of terminal models that closely approximates the LDGP. However, model selection still yields advantages over imposing a given theory model or simple model averaging when implementing impulse- indicator saturation (IIS) within the selection procedure. IIS includes a set of saturating impulse dummies in the GUM, see Hendry, Johansen and Santos (2008) [41], Johansen and Nielsen (2009, 2016) [52, 53], and Hendry and Santos (2010) [46], which acts as a robust method when there is model misspecification by accounting for location shifts and outliers in omitted relevant variables, helping to mitigate the adverse impacts of induced location shifts on intercepts and equation standard errors.† Interestingly, location shifts in omitted variables do not affect slope parameters, even when correlated with included variables, so thick modelling would be still be valid on the slope parameters in the case where the GUM is underspecified and the omitted variables shift. This suggests that IIS should always be implemented to provide robustness against misspecified GUMs and to account for outliers and location shifts in-sample. In dynamic models, locating the precise timing of shifts can be difficult which may result in multiple terminal models in which the specifications differ due to the timing of retained impulses. Thick modelling would average over these terminal models, capturing the uncertainty of the timing of possible location shifts. In many cases, outliers or shifts are due to known reasons such as a policy change in a given month/quarter. Institutional knowledge of this kind would allow the researcher to narrow down the set of terminal models to retain those in which impulses correspond to known events. However, if there are dynamic effects which make policy effects slow to feed through, or the researcher is uncertain of such events leading to outliers or shifts in the data, then thick modelling provides a way of capturing the timing uncertainty. Granger’s proposal of thick modelling in which alternative specifications of similar quality are combined rather than selecting one and discarding all others is made feasible by using Gets which results in a set of terminal models that satisfy his criterion of similar valid models. Granger’s pragmatic approach to modelling enables model averaging and model selection to complement each other. In the next section we briefly review the literature, before demonstrating the approach using artificial data. 3. Literature on thick modelling Granger’s proposal for thick modelling was motivated by his work on forecasting, where pooling of forecasts often outperformed the selected ‘best’ forecast. The subsequent litera- ture using thick modelling has followed in this tradition, focussing on forecasting applica- tions, see for example McNelis and McAdam (2004) [56] and Albacete and Espasa (2005) [3] for inflation forecasting examples, and Aiolfi and Favero (2005) [1] for an application †Robustness in the statistics literature refers to methods that perform well under non-standard distribu- tions, and in particular when there are large-sized observations or outliers. Here we use a broader definition of robustness, whereby methods are robust if they have good properties against many forms of misspecifica- tion, including outliers, location shifts, omitted variables, incorrect distributional shape, non-stationarity, misspecified dynamics and non-linearities. J. Castle / Eur. J. Pure Appl. Math, 10 (1) (2017), 133-156 142 yt 0 10 20 30 40 50 60 70 80 90 100 -5.0 -2.5 0.0 2.5 5.0 7.5 10.0 12.5 yt Figure 2: Dependent variable for the artificial data set to the predictability of stock returns. Pesaran and Timmermann (2004) (see [62]) suggest thick modelling as a technique for real-time forecasting. Granger’s second motivation for thick modelling came from his work on portfolio theory, and thick modelling has been applied in the multivariate volatility literature by papers such as Amendola and Storti (2015) [4], and Pesaran, Schleicher and Zaffaroni (2009) [61], who use thick modelling as a way to address model uncertainty in multivariate volatility models. The literature has mainly focussed on the use of thick modelling in forecasting, where forecast combination has a long pedigree, but the proposal in Granger and Jeon (2004) [30] was not specific to forecasting, and thick modelling can be applied in-sample successfully as well, as we now explore. 4. Applying Gets with thick modelling to a data set In this section, we apply Autometrics selection to a single draw from a known LDGP and then use thick modelling over the resulting terminal models. Assume the LDGP is an I(0) autoregressive distributed lag model which contains two location shifts within the sample that shifts the unconditional mean but not the slope parameters: yt = β0 + βyyt−1 + 10∑ i=1 βixi,t + δ1(tL T , but selection is perfectly feasible using Autometrics within a few seconds, see Castle and Hendry (2014) [13]. Applying selection at the tighter target significance level of 2.5% results in two terminal models being retained, see table 2. We would expect approximately 5 irrelevant variables to be retained on average at this significance level but this is not the case for the given draw. At 10% 7 terminal models are retained, with the only J. Castle / Eur. J. Pure Appl. Math, 10 (1) (2017), 133-156 149 difference in terminals due to the impulse and step indicators. At this loose significance level Autometrics heavily overfits, retaining between 68 and 74 indicator variables due to the downward bias in the equation standard error by removing many observations. Eliminating sequentially marginally significant indicators results in the significance of other indicators collapsing, which we term the ‘house of cards’ problem which can be present when the selection significance level is too loose. At 1% a unique terminal model is retained so no thick modelling is required. TM1 TM2 Union Pooled Constant 0.089 0.126 0.081 0.099 yt−1 0.640 0.666 0.641 0.649 x1,t 0.357 0.342 0.354 0.351 x2,t−1 0.428 0.415 0.423 0.422 x3,t 0.302 0.275 0.297 0.291 x4,t 0.512 0.477 0.503 0.497 x5,t 0.766 0.757 0.761 0.761 1t≤51 -2.528 -2.324 -2.547 -2.426∗ 1t≤82 . 2.338 0.611 1.169∗ 1t≤83 2.555 . 1.967 1.278∗ LL -149.0 -150.3 -148.9 σ̂ 1.126 1.140 1.131 AR(2) 0.230 0.158 0.229 ARCH(1) 0.118 0.254 0.118 Normality 0.553 0.470 0.557 Hetero 0.141 0.146 0.132 Chow(70%) 0.718 0.623 0.745 Table 2: Terminal models from Autometrics undertaking selection at a target size of 2.5% for GUM (9). Coefficient estimates reported, with the unweighted average of TM1, TM2 and Union in the Pooled column. p-values for diagnostic tests reported, along with the log-likelihood (LL) and estimated equation standard error (σ̂). ∗ denotes averaging across TM1 and TM2 but excluding the union model. The benefits of thick modelling are most apparent in this well-specified case where the exogenous regressors are consistently selected across terminals and the only difference between the two terminal models highlights the difficulty in pinpointing the precise date of the structural break. Both terminals are mutually encompassing so thick modelling reflects the uncertainty of the break timing. Thick modelling proposes to average the coefficient estimates if a point estimate is required, although just reporting the range would be in keeping with the thick modelling philosophy. Mean coefficient estimates for step indicators should be treated carefully if averaging with the union of the terminal models. TM1 and TM2 retain different break dates, but the union combines these resulting in a joint effect of the step shift given by the sum of the coefficients in the union model. Therefore we pool the step shift coefficients by excluding the union model. In order to judge how well the thick modelling performs we should evaluate the pooled model against the estimated LDGP. Even if a researcher knew the LDGP specification J. Castle / Eur. J. Pure Appl. Math, 10 (1) (2017), 133-156 150 yt E[y |x] Min E[y |x] Min 0 20 40 60 80 100 0 10 yt E[y |x] Min E[y |x] Min yt ŷT+h |T+h−1 MIN ŷT+h |T+h−1 MAX ŷT+h |T+h−1 MEAN 95 100 105 110 115 120 -5 0 yt ŷT+h |T+h−1 MIN ŷT+h |T+h−1 MAX ŷT+h |T+h−1 MEAN Figure 7: In-sample model fit and forecasts for selection from congruent GUM she would still need to estimate the model, and as such could make mistakes based on inference from the estimated LDGP. The estimated LDGP is given in equation (10): yt = 0.704 (0.041) yt−1 + 0.265 (0.087) x1,t + 0.313 (0.086) x2,t + 0.289 (0.093) x3,t + 0.441 (0.098) x4,t +0.732 (0.113) x5,t + 2.141 (0.400) 1{50